1
00:00:00,000 --> 00:00:05,000
Change the view. Keep the dynamics.
Complete video in. An evolving scene out.

2
00:00:05,000 --> 00:00:08,300
Geometry, illumination and optical media
jointly shape the recorded appearance.

3
00:00:08,300 --> 00:00:12,000
Explicitly modeling all their changes is difficult.
The full video records their joint appearance over time.

4
00:00:12,000 --> 00:00:16,000
A single RGB frame observes one instant.
GeoV2V conditions on the complete source clip.

5
00:00:16,000 --> 00:00:20,000
Source and target latents share video self-attention,
without an added frame-isolation mask.

6
00:00:20,000 --> 00:00:24,200
The selected instants illustrate access across time.
Source conditioning follows ReCamMaster.

7
00:00:24,200 --> 00:00:30,700
Now move the camera.
Measured LiDAR supplies the spatial anchor.

8
00:00:30,700 --> 00:00:34,500
The same scan opens into target views
at 1, 2 and 4 metre lateral offsets.

9
00:00:34,500 --> 00:00:37,467
Original generated outputs appear beneath
corresponding calibrated projections.

10
00:00:37,467 --> 00:00:40,800
Training pairs come from Para4D in CARLA:
varied maps, building assets and dynamic traffic.

11
00:00:40,800 --> 00:00:43,700
Seven parallel positions, each with five camera views,
move together and capture synchronized videos.

12
00:00:43,700 --> 00:00:46,633
Source video, geometry and camera poses condition prediction.
The captured target video provides paired supervision.

13
00:00:46,633 --> 00:00:50,467
Three dynamic scenes, seven lateral positions each.
The marked source and ground truth stay synchronized.

14
00:00:50,467 --> 00:00:55,000
The training diagram unfolds into the model.
Two target-view priors guide video generation.

15
00:00:55,000 --> 00:00:59,000
Dense geometry controls DiT blocks 1–15.
LiDAR controls blocks 16–30.

16
00:00:59,000 --> 00:01:03,000
Full source latents join noisy target latents.
Decode the target sequence into a new-view video.

17
00:01:03,000 --> 00:01:08,000
The source and all three generated views
advance on the same capture clock.

18
00:01:08,000 --> 00:01:11,200
Temporal evidence supports the scene's dynamics
through a change of viewpoint.

19
00:01:11,200 --> 00:01:15,000
A turning signal is an event through time.
Follow the marked light in the published frames.

20
00:01:15,000 --> 00:01:18,500
The event is temporal.
So is the conditioning.

21
00:01:18,500 --> 00:01:24,000
The supplementary rain example also shows
changing illumination and optical effects.

22
00:01:24,000 --> 00:01:29,000
GeoV2V: full video context, geometry-grounded generation.
Change the view. Keep the dynamics.
