ECCV 2026 ↗

GeoV2V

Geometry-grounded video diffusion for driving scene generation.

THE RESEARCH FILM · 01:29 · 1080P · EN / 中文
Compare the views ↓
Chapters & downloads

One drive. A different lane.

1.5× synchronized replay
RECORDED VIDEO 0 m GEOV2V 2 m
Source
Generated view
Move the viewpoint
00.0SAME CAPTURE TIME

Actual model outputs · original frames · 1.5× synchronized replay · precomputed 1 / 2 / 4 m lateral shifts

Full video as source contextGeometry for viewpoint controlDynamics across generated frames

Source footage and generated views share one capture clock. The film links the same scene to measured geometry; its reconstructed projections illustrate the process. Scene provenance ↗

GeoV2V: Geometry-Grounded Video Diffusion Model for Driving Scene Generation

1 Global Institute of Future Technology, SJTU2 Global College, SJTU3 School of Mechanical Engineering, SJTU4 Shanghai Pudong Development Bank

* Equal contribution   ·   † Corresponding author   ·   ‡ Work done at GIFT, SJTU

WHY THE WHOLE VIDEO MATTERS

The camera sees
more than geometry.

Moving objects. Blinking lights. Rain and lens droplets. Explicitly modelling every changing factor is difficult. A driving video already records their combined appearance through time.

RECORDED INPUT

Captured video

The observed scene, through time.

Optical medium

Rain, fog and lens droplets.

Illumination

Signals, brake lights and ambient light.

Geometry

Moving agents and static structure.

Conceptual image formation from Fig. 2
FROM FORMATION TO GENERATION

Geometry anchors the view.
Video carries the appearance.

The paper’s source / left-2 m pairs retain changing turn and brake lights. Full-video context supplies temporal evidence alongside geometric guidance.

See how the full clip is used

Schematic layers and moving rays explain the paper’s image-formation argument; they are not a physical rendering simulation. The paired frames above are an unchanged excerpt of Fig. 2. Dynamic-effect preservation is demonstrated qualitatively.

A scene is more
than one moment.

A single RGB frame cannot reveal how a turn signal blinks or a vehicle moves. GeoV2V conditions on the full source clip, so the model can use appearance as it changes over time.

A

Single-frame conditioning

One source frameOne RGB observation

Appearance at one instant.
No observed source-video evolution.

B

GeoV2V · full-video context

CONNECTED THROUGH TIME
SOURCE CLIP t₁ → tₙ
Source video example at sample 1t1
Source video example at sample 2t2
Source video example at sample 3t3
Source video example at sample 4t4
Shared video attentionSource and target latents, connected
Generated video example at sample 1t1
Generated video example at sample 2t2
Generated video example at sample 3t3
Generated video example at sample 4t4
GENERATED CLIP New viewpoint · evolving scene

One complete source clip. No added frame-isolation mask between source and target video latents.

Conceptual comparison with single-frame RGB conditioning; four displayed frames illustrate a full clip. Video latents are VAE-encoded. Full-video conditioning follows ReCamMaster; GeoV2V combines it with complementary geometric controls. Curves illustrate accessible context, not measured attention weights.

Keep the world.
Choose the view.

The source video supplies temporal context. Two geometric priors tell the model where the scene belongs in the requested camera view.

01

Measured precision.

Sparse LiDAR preserves reliable 3D structure and is reprojected into the target view.

LiDAR → ConvNeXt → blocks 16–30
02

Broader coverage.

Depth-derived geometry and depth-aware filling provide a complementary dense prior.

Dense prior → VAE + mask → blocks 1–15
Geometry preprocessing · original paper figure

Shared context.
Complementary control.

Clean source-video latents and noisy target-video latents enter a shared transformer. Two geometry controllers guide the target sequence at different depths.

Full source clip VAE-encoded
Noisy target clip + requested cameras
Temporal concatenation ↓
Dense geometryBlocks 1–15
LiDAR geometryBlocks 16–30
Shared self-attention across the video latents↓Decode the generated video
Wan 2.1 · 1.3B backbone · two 61M controllers

The scene stays alive.

From moving vehicles to rain on the lens. Explore the original, labeled supplementary comparisons.

Novel-view generation

View control, measured.

Reported numbers from the paper.
Explore the setting before comparing the scores.

FVD ↓
232.06GEOV2V · FVD AT 2 M

Lowest FVD among the compared methods at every tested real-world lateral offset: 1, 2, and 4 metres.

Paper Table 2 · off-record trajectories

What each condition contributes +

Remove a condition. Measure the change.

Paper Table 4 measures component removal at a 2 m lateral shift on Para4D. Removing dense geometry reduces PSNR from 23.88 to 17.18; replacing synchronized Para4D supervision with misaligned side views raises LPIPS from 0.051 to 0.102.

Component ablation · Para4D · 2 m lateral shift · best scores underlined
Configuration PSNR ↑ SSIM ↑ LPIPS ↓

The video-conditioned model changes these aggregate metrics modestly; the paper also discusses qualitative detail preservation. The original w/o Para4D row replaces synchronized supervision with misaligned side views from Para4D and Waymo.

Seven positions.
One shared instant.

Synchronized multi-view supervision in CARLA. Seven lateral positions, five cameras at each position, and dynamic driving scenes.

Parallel trajectories capture the same dynamic scene at the same time. These CARLA recordings provide paired source and target videos.

Offsets: −4, −2, −1, 0, +1, +2, +4 mPaired ground truth35 cameras
Scope and limitations +

Large disocclusions and inaccurate depth can weaken geometric guidance. Video-VAE artifacts also remain a limitation. The result selector plays precomputed model outputs; reconstructed projections in the film illustrate geometric conditioning.

Read supplementary material ↗

Paper. Film. Evidence.

Everything behind this project, in one place.

Explore the original figures 19 panels · main paper & supplement+

Original research images, with their method labels and annotations retained. Download the media inventory ↓

Complete original supplementary videoAll original sections · web encode · play here
Original research material, with its submitted sections and labels.Download MP4 ↓

Model implementation, checkpoints, and the full Para4D dataset are not included in this project-page repository. Release links will be added when available.

Cite GeoV2V

@InProceedings{xi2026geov2v,
  author    = {Xi, Yuchen and Guo, Yuhan and Hou, Chenwei and
               Liu, Zixu and Fang, Jin and Liu, Jason and Yang, Ruigang},
  title     = {GeoV2V: Geometry-Grounded Video Diffusion Model
               for Driving Scene Generation},
  booktitle = {Computer Vision -- ECCV 2026},
  year      = {2026},
  series    = {Lecture Notes in Computer Science},
  volume    = {17012},
  pages     = {656--673},
  publisher = {Springer Nature Switzerland},
  doi       = {10.1007/978-3-032-37044-0_36}
}

The Waymo driving sequences and measured geometry use the Waymo Open Dataset, provided by Waymo LLC under its License Agreement for Non-Commercial Use. Access and use are governed by those terms. Data and code attribution ↗

Original paper figure. Open full resolution to inspect details.

Open full resolution ↗