LingBot-Map

Streaming 3D Reconstruction with Geometric Context Transformer

Method

Streaming 3D reconstruction is fundamentally a question of memorywhat to keep, and in what form. LingBot-Map answers this with Geometric Context Attention (GCA), a small but structured streaming state that is learned endtoend. GCA maintains three complementary contexts: an anchor for coordinate and scale grounding, a local pose‑reference window for dense local geometry, and a trajectory memory that compresses the full history into compact perframe tokens keeping memory and compute per frame nearly constant on sequences of 10,000+ frames at ~20 FPS.

Pipeline of LingBot-Map
Pipeline of LingBot-Map. A DINO backbone extracts image features, which are refined through alternating layers of Frame Attention and GCA. Within GCA, the current view aggregates information from the Anchor Context, the local Pose‑Reference Window, and the Trajectory Memory. Task‑specific heads then predict camera pose and depth maps.

Demo

Select a scene below — the large viewer plays the corresponding streaming point‑cloud reconstruction.

Multi-Room Traversal

Camera Trajectory Estimation

Streaming pose estimates versus ground truth across diverse benchmarks.

Tanks & Temples · Barn
Oxford Spires · Keble College
Oxford Spires · Observatory Quarter