Chapters & downloads
One drive. A different lane.
Actual model outputs · original frames · 1.5× synchronized replay · precomputed 1 / 2 / 4 m lateral shifts
Source footage and generated views share one capture clock. The film links the same scene to measured geometry; its reconstructed projections illustrate the process. Scene provenance ↗
GeoV2V: Geometry-Grounded Video Diffusion Model for Driving Scene Generation
1 Global Institute of Future Technology, SJTU2 Global College, SJTU3 School of Mechanical Engineering, SJTU4 Shanghai Pudong Development Bank
The camera sees
more than geometry.
Moving objects. Blinking lights. Rain and lens droplets. Explicitly modelling every changing factor is difficult. A driving video already records their combined appearance through time.
Captured video
The observed scene, through time.
Optical medium
Rain, fog and lens droplets.
Illumination
Signals, brake lights and ambient light.
Geometry
Moving agents and static structure.
Geometry anchors the view.
Video carries the appearance.
The paper’s source / left-2 m pairs retain changing turn and brake lights. Full-video context supplies temporal evidence alongside geometric guidance.
See how the full clip is usedSchematic layers and moving rays explain the paper’s image-formation argument; they are not a physical rendering simulation. The paired frames above are an unchanged excerpt of Fig. 2. Dynamic-effect preservation is demonstrated qualitatively.
A scene is more
than one moment.
A single RGB frame cannot reveal how a turn signal blinks or a vehicle moves. GeoV2V conditions on the full source clip, so the model can use appearance as it changes over time.
Single-frame conditioning
One RGB observation
Appearance at one instant.
No observed source-video
evolution.
GeoV2V · full-video context
CONNECTED THROUGH TIME
t1
t2
t3
t4
t1
t2
t3
t4
One complete source clip. No added frame-isolation mask between source and target video latents.
Reduced motion · static illustrationConceptual comparison with single-frame RGB conditioning; four displayed frames illustrate a full clip. Video latents are VAE-encoded. Full-video conditioning follows ReCamMaster; GeoV2V combines it with complementary geometric controls. Curves illustrate accessible context, not measured attention weights.
Keep the world.
Choose the view.
The source video supplies temporal context. Two geometric priors tell the model where the scene belongs in the requested camera view.
Measured precision.
Sparse LiDAR preserves reliable 3D structure and is reprojected into the target view.
LiDAR → ConvNeXt → blocks 16–30Broader coverage.
Depth-derived geometry and depth-aware filling provide a complementary dense prior.
Dense prior → VAE + mask → blocks 1–15Shared context.
Complementary control.
Clean source-video latents and noisy target-video latents enter a shared transformer. Two geometry controllers guide the target sequence at different depths.
The scene stays alive.
From moving vehicles to rain on the lens. Explore the original, labeled supplementary comparisons.
Novel-view generation
View control, measured.
Reported numbers from the paper.
Explore the setting before
comparing the scores.
Lowest FVD among the compared methods at every tested real-world lateral offset: 1, 2, and 4 metres.
Paper Table 2 · off-record trajectoriesWhat each condition contributes +
Remove a condition. Measure the change.
Paper Table 4 measures component removal at a 2 m lateral shift on Para4D. Removing dense geometry reduces PSNR from 23.88 to 17.18; replacing synchronized Para4D supervision with misaligned side views raises LPIPS from 0.051 to 0.102.
| Configuration | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|
The video-conditioned model changes these aggregate metrics modestly; the paper also discusses qualitative detail preservation. The original w/o Para4D row replaces synchronized supervision with misaligned side views from Para4D and Waymo.
Seven positions.
One shared instant.
Synchronized multi-view supervision in CARLA. Seven lateral positions, five cameras at each position, and dynamic driving scenes.
Parallel trajectories capture the same dynamic scene at the same time. These CARLA recordings provide paired source and target videos.
Scope and limitations +
Large disocclusions and inaccurate depth can weaken geometric guidance. Video-VAE artifacts also remain a limitation. The result selector plays precomputed model outputs; reconstructed projections in the film illustrate geometric conditioning.
Read supplementary material ↗Paper. Film. Evidence.
Everything behind this project, in one place.
Explore the original figures 19 panels · main paper & supplement+
Original research images, with their method labels and annotations retained. Download the media inventory ↓
Complete original supplementary videoAll original sections · web encode · play here
Model implementation, checkpoints, and the full Para4D dataset are not included in this project-page repository. Release links will be added when available.
Cite GeoV2V
@InProceedings{xi2026geov2v,
author = {Xi, Yuchen and Guo, Yuhan and Hou, Chenwei and
Liu, Zixu and Fang, Jin and Liu, Jason and Yang, Ruigang},
title = {GeoV2V: Geometry-Grounded Video Diffusion Model
for Driving Scene Generation},
booktitle = {Computer Vision -- ECCV 2026},
year = {2026},
series = {Lecture Notes in Computer Science},
volume = {17012},
pages = {656--673},
publisher = {Springer Nature Switzerland},
doi = {10.1007/978-3-032-37044-0_36}
}