# Camera-ready media coverage

This inventory maps the active image references in the accepted camera-ready main paper and supplementary LaTeX to public website assets. The authoritative source directory is `最新论文/overleaf-69a4f399287924273ceecf8e`; supplementary video sources are in `418/Supplementary_Material`. Original results remain original results: the exports add no model inference, synthetic intermediate prediction, or retouching.

All **19 active image sources** are represented: nine from the main paper, including both panels of its first figure, and ten from the supplementary material. Commented-out LaTeX figures are excluded. `assets/manifest.json` records source paths, export paths and dimensions; the eight newly added figures also carry source/export SHA-256 values. Their full-resolution raster exports preserve the original decoded RGB pixels as lossless WebP. The PDF figure is rendered with Poppler at 3400 pixels wide and its original PDF is available unchanged. Separate `-preview.webp` thumbnails use at most 960 pixels on either edge and are included for lightweight reuse; the figure library opens full-resolution originals for detailed inspection.

| Camera-ready source | Public asset in `assets/images/` | Evidence or scope |
| --- | --- | --- |
| `src_cr/overview.pdf` | `overview.webp`, `overview.pdf` | Overview of the approach |
| `src_new/Bullet_ECCV_sequence.jpg` | `dynamic-lighting-sequence.webp` | Time-indexed source and generated driving frames |
| `src_cr/motivation.jpg` | `motivation.webp` | Geometry, illumination and optical media in image formation; temporal examples |
| `src_new/Para4D.png` | `para4d.webp` | CARLA scene and synchronized parallel-camera construction |
| `src_new/Geo_Preprocess_final_cropped.pdf` | `geometry-priors.webp`, `geometry-priors.pdf` | Sparse LiDAR and dense geometry preprocessing |
| `src_cr/pipeline.pdf` | `pipeline.webp`, `pipeline.pdf` | Actual model architecture and conditioning paths |
| `src_cr/ECCV-2026-bullet_compressed_3000w.jpg` | `qualitative-comparison.webp` | Original labeled method comparisons |
| `src_new/light.jpg` | `lighting-comparison.webp` | Dynamic-lighting comparisons |
| `src/abl.drawio.png` | `ablation.webp` | Original visual ablations |
| `src_supp/more_results.jpg` | `more-waymo-results.webp` | Further Waymo results at lateral shifts of 1, 2 and 4 m |
| `src_supp/Shift_Rotation.jpg` | `translation-rotation.webp` | Qualitative zero-shot translation and rotation, including 2 m shifts and 15-degree rotations |
| `src/waymo_interp.png` | `waymo-interpolation.webp` | Waymo frame interpolation, with original baseline labels |
| `src/nuscenes_interp_2.png` | `nuscenes-interpolation.webp` | nuScenes frame interpolation, with original baseline labels |
| `src_supp/carla_nvs_1.pdf` | `para4d-qualitative.webp`, `para4d-qualitative.pdf` | Synthetic Para4D qualitative novel-view results |
| `src_supp/xld.png` | `xld-static-scenes.webp` | Original supplementary illustration of an earlier static-scene dataset |
| `src_supp/carla_multi_view.pdf` | `para4d-multiview.webp`, `para4d-multiview.pdf` | Synchronized five-camera observations |
| `src_supp/carla_multi_shifted.pdf` | `para4d-lane-supervision.webp`, `para4d-lane-supervision.pdf` | Synchronized cross-lane training supervision |
| `src_supp/limitation.png` | `limitations.webp` | Domain-gap example: plausible textures with incorrect colors under occlusion |
| `src_supp/vae_fault.png` | `vae-temporal-artifacts.webp` | Separate temporal-VAE failure example, comparing frames `4n` and `4n+1` |

## Original supplementary videos and excerpt boundaries

There are five independent supplementary source videos and two combined exports of those materials in the source folder. The webpage offers convenient excerpts and the full combined supplementary video. The newly authored explainer is a separate asset, `assets/video/geov2v-explainer.mp4`.

Times below are relative to each named independent source. Ranges describe the requested extraction interval; the manifest is the machine-readable record. Existing excerpts are intentionally retained so that older editorial timelines and crop coordinates remain reproducible.

| Independent source | Source duration | Website asset and coverage |
| --- | --- | --- |
| `GeoV2V_Main_Comparison.mp4` | 19.73 s | `comparison-reconstruction.mp4`: 2.1-9.7 s; `comparison-generation.mp4`: 12.0-19.5 s. Original baseline labels and 1/2/4 m columns are retained. `input-0m`, `result-1m`, `result-2m`, `result-4m` and `hero` are crops of the first excerpt. |
| `GeoV2V_Main_Results.mp4` | 21.03 s | `waymo-overview.mp4`: 0-10 s; `rain-triptych.mp4`: 10.6-20.6 s. The latter also supplies the three individually cropped rain videos. |
| `GeoV2V_Tested_On_Para4D.mp4` | 8.10 s | `para4d-results.mp4`: full 0-8.1 s. |
| `Para4D_Dataset_Visualization_Simple.mp4` | 484 frames / 16.1333 s | Existing `para4d-lanes.mp4`: 0-8.0 s, seven lateral offsets across three scenes. New `para4d-directions.mp4`: source frames 243-483, starting at the exact cut at 8.1 s; all 241 remaining frames, 8.0333 s, five viewing directions across three scenes. The 0.1 s between the two excerpt ranges remains available in the full compilation. |
| `Para4D_Dataset_Visualization_Full_Cameras.mp4` | 25.13 s | `para4d-rigs.mp4`: 0-25.1 s, the full seven-position by five-camera presentation apart from the final approximately 0.03 s. The full compilation retains the source ending. |

### Full combined supplementary video

`assets/video/original-supplement.mp4` is a web encode of `ECCV2026_GeoV2V_Supplementary_Video_YouTube_1440p_150MB.mp4`. It keeps the full original content in this order: **Main Comparison → Main Results → Tested On Para4D → Simple dataset visualization → Full Cameras dataset visualization**. This order and all five kinds of material were visually checked in the source compilation at 4, 14, 25, 35, 47, 54, 64, 77 and 86 seconds. The much larger 1 GB combined export is not copied into the public repository.

- Full source and web export: **2704 video frames, 30 fps, 90.1333 s, 2496 × 1440, unchanged aspect ratio**.
- Encoding: H.264 `libx264`, preset `medium`, CRF 22, maximum rate 6 Mbps, buffer 12 Mbps, `yuv420p`, MP4 fast start. No trimming, scaling, frame interpolation or replacement.
- Original AAC audio packets are copied; their source/export packet MD5 values match. Audio is retained even if the original track is effectively silent.
- Export size: **59,680,786 bytes**, below the 100 MB per-file limit.
- The actual source and export SHA-256 values, poster path, frame count, encoding settings and validation record are stored in `assets/manifest.json`.

The new camera-direction excerpt preserves native **1690 × 720, 30 fps** and all remaining 241 video frames. It uses H.264 CRF 19, retains the source framing, and encodes the remaining original audio as AAC. Its exact extraction is video `trim=start_frame=243,setpts=PTS-STARTPTS`, audio `atrim=start=8.1,asetpts=PTS-STARTPTS`.

## Scientific interpretation and verification

- Image formation in the motivation figure is conceptual. It illustrates geometry, illumination and optical media affecting recorded appearance. The GeoV2V architecture does not explicitly simulate physical light transport.
- Full-video conditioning consumes the supplied source clip through a video VAE. This is not a claim of mathematically lossless encoding or guaranteed preservation of every detail.
- The paper's real-world off-record lane-shift examples do not supply target-view ground truth. Original figure labels are preserved; a recorded input column labeled `GT` must not be reinterpreted as ground truth for a shifted target camera.
- Frame-interpolation figures are identified as interpolation, separately from lateral-extrapolation evidence. Translation/rotation panels are qualitative examples, not additional quantitative benchmark results.
- `limitations.webp` and `vae-temporal-artifacts.webp` are different failure cases. The first concerns domain gap and color uncertainty; the second concerns temporal VAE artifacts. See supplementary LaTeX lines 177-198 for the original captions and qualifications.
- Seven new raster figure exports were verified for decoded RGB pixel identity with their source images. The PDF rendering and all eight thumbnail compositions were visually inspected for preserved content, labels and aspect ratio.
- Both new MP4 exports completed a full audio/video decode without errors. Source/export frame counts were checked. All 2704 frames are retained in the full compilation; all 241 frames from source index 243 onward are retained in the new direction excerpt. Original combined-video audio packet hashes match.

Website figure annotations and film crops may guide attention to source evidence. They do not create additional experimental results. Source cropping or timing for editorial animation should remain separately documented in the film production provenance.

The main film's Para4D pair expands into the original Simple video's **three-scene by seven-offset** montage. Its columns are left 4 m, left 2 m, left 1 m, 0 m, right 1 m, right 2 m and right 4 m. The selected source and supervised target are the top-row fourth and second cells. Synchronization applies within each scene row. This montage uses the first take before the layout change at 8.1 seconds; that boundary changes to five viewing directions and restarts scene playback, so it is not used as a continuous extension of the seven-offset capture.
