StreamRig: Exploiting Intra-Rig Geometry
for Streaming Multi-Camera Odometry

Yufei Wei1Shuhao Ye1Qi Wang2Xin Zheng1Qing Huang1Rong Xiong1Yue Wang1,∗

1 Zhejiang University2 South China University of Technology

∗ Corresponding author

Joint rig geometry. Compact causal state. Continuous odometry.
Frozen multi-view geometry becomes causal rig odometry, with only 74.6M trainable parameters.

Click the map to re-anchor
Loading point cloud…
Five-camera map
Drag to rotate · Right-drag to pan · Scroll to zoom
Height above ground0–12 m
Current rig
Mapping cameras
Pose source
Point color
Frame 381 / 501
StreamRig trajectory Current rig & camera frusta
Show synchronized input images +

Abstract

Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present StreamRig, a freeze-and-stream framework that builds causal streaming odometry for calibrated rigs on a frozen multi-view 3D foundation model.

The frozen front-end jointly perceives synchronized views using rig calibration. A Rig-Resampler compresses their features, a CausalBridge applies causal attention with a key-value cache, and a lightweight head regresses rig poses. Periodic re-anchoring supports long sequences. Only these modules are trained—74.6M parameters—with relative poses as the sole supervision.

Two-stage training transfers group relocalization ability to causal rig odometry. We evaluate on NCLT, TartanGround, KITTI-360, and the ZJH humanoid-robot dataset, where real-world evaluation uses simulation-trained weights. Across all four datasets, StreamRig achieves lower translation and rotation drift than the evaluated non-oracle monocular streaming and rig-aware offline models.

Method

Perceive the rig jointly, compress its geometry, and carry it through a causal stream.

StreamRig architecture, causal attention, two-stage training, and re-anchoring
StreamRig consists of a frozen multi-view front-end, a Rig-Resampler, a CausalBridge, and a pose head. The trainable modules first learn group relocalization, then learn causal motion from rig windows.

Joint geometry

Every arriving rig is encoded once, with camera calibration supplied to the frozen multi-view model.

Compact causal state

Each camera contributes 16 latent tokens. Each query interacts with its own anchor snapshot and cached history.

Periodic re-anchoring

The last rig becomes the next anchor. Relative poses compose into a continuous trajectory with bounded active state.

Experimental Results

Four rigs spanning campus navigation, simulation, urban driving, and a humanoid robot.

Example trajectories of StreamRig, G2G-chain, CUT3R and LingBot-Map against ground truth on NCLT, TartanGround, KITTI-360 and ZJH
Example trajectories on the four rigs.

Streaming odometry on four rigs

Table I. t_rel in percent, r_rel in degrees per 100 m, ATE in meters. Lower is better.
MethodCamerasNCLT5-cam surround · 2 traj. / 12.29 kmTartanGround4×90° ring · 10 scenes / 12.60 kmKITTI-360stereo + 2 fisheye · 2 scenes / 13.63 kmZJH (real robot)stereo + 2 oblique · 3 scenes / 0.36 km
trelrrelATEtrelrrelATEtrelrrelATEtrelrrelATE

best second best · trel (%) · rrel (°/100 m) · ATE (m) · lower is better · grey: oracle with ground-truth relative pose

Ablations and analysis

Ablations, one factor at a time
GroupVariantNCLTTartanGroundZJH (real)
trelrreltrelrreltrelrrel

Design choices matter. Mixed windows, displacement-normalized translation, and the per-query anchor snapshot all reduce drift; removing group-relocalization pre-training increases translation drift by 4.6–14.9×.

Mono, stereo, and full rig on KITTI-360
Camerastrel ↓rrel ↓SE3 ↓Sim3 ↓

More cameras, less drift. The stereo pair adds depth, and the two side cameras widen the view; the full rig is best on every metric.

Frozen front-end on NCLT
Front-endtrel (%) at NBest
6122448trelrrel

A swappable front-end. Depth Anything 3 and π³X also work as frozen front-ends; MapAnything, which conditions on camera rays, exposes rig geometry most directly.

Re-anchoring distance and training horizon
Translation and rotation drift versus re-anchoring interval for three training horizons on NCLT, with training costs
Translation (filled) and rotation (open) drift on NCLT versus re-anchoring interval, for three training horizons.

Re-anchoring sets the operating point. Drift is lowest at intermediate intervals; longer training windows extend the usable range at higher cost.

Compute and memory budget per arrival
MethodInputms ↓Thr. ↑Mem ↓

A whole rig for less than one monocular frame. StreamRig processes five cameras in 26.2 ms with 2.6 GiB, faster than the fastest monocular streamer on a single frame.

Video

The method and synchronized trajectory comparisons across all four datasets.

Citation

@misc{wei2026streamrigexploitingintrariggeometry,
      title={StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry},
      author={Yufei Wei and Shuhao Ye and Qi Wang and Xin Zheng and Qing Huang and Rong Xiong and Yue Wang},
      year={2026},
      eprint={2609.40244},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.40244},
}