Hierarchical geometry fusion
Geometry from three VGGT depths is aligned with successive language-decoder stages. Each branch uses normalization, 2 × 2 spatial grouping, a depth-specific projection, a token-wise gate, and a learned scale.
Spatial reasoning · Streaming navigation

Selective Geometry Across Representation Depth and Navigation Time for Vision-Language Navigation
What geometry should a navigation agent use?
And what should it remember as it moves?
01 / Overview
Vision-language navigation requires an embodied agent to align language with visual observations while maintaining a spatially coherent understanding of its environment. Geometry foundation models expose representations throughout their processing hierarchy, but using only a terminal feature can leave complementary geometric information inaccessible to the navigation policy.
AdaGeoVLN selects geometry across both representation depth and navigation time. Hierarchical GFM–VLM fusion couples earlier, intermediate, and later VGGT representations to successive stages of the navigation policy. Navigation-aware memory retains historical VGGT global-attention key–value states according to instruction relevance, geometric confidence, and transition novelty.
Experiments on R2R-CE and RxR-CE show strong performance with a single RGB stream and no additional navigation-specific external training data. Controlled ablations support the value of multi-depth fusion, while navigation-aware retention improves SR and SPL at a similar measured KV footprint and reduces memory relative to larger-memory temporal retention.
02 / Method
The policy combines a Qwen3.5-4B vision-language model with a frozen VGGT-1B geometry encoder. Two mechanisms determine how geometry enters reasoning and remains available at future steps.
Geometry from three VGGT depths is aligned with successive language-decoder stages. Each branch uses normalization, 2 × 2 spatial grouping, a depth-specific projection, a token-wise gate, and a learned scale.
Historical and current VGGT global-attention KV states compete for a fixed per-layer budget. Three complementary signals determine which patch tokens to retain:
Maximum centered cosine similarity between projected shallow tokens and instruction segments; cached at insertion.
The minimum of calibrated depth and point-map confidence, requiring both predictors to be reliable.
Nearest-other-frame descriptor novelty across all observed frames, refreshed at each step. Relevance and confidence enter the score separately.
03 / Benchmark results
Results on 1,839 R2R-CE and 3,669 RxR-CE Val-Unseen episodes. The tables preserve the paper’s sensing assumptions and external training data so the comparisons can be read in context.
↑ Higher is better ↓ Lower is better
| Method | Observation | NE ↓ | OS ↑ | SR ↑ | SPL ↑ | External data |
|---|---|---|---|---|---|---|
| HPN+DN | Pano. + odo. + depth | 6.31 | 40.0 | 36.0 | 34.0 | — |
| Sim2Sim | Pano. + odo. + depth | 6.07 | 52.0 | 43.0 | 36.0 | — |
| VLN⟳BERT | Pano. + odo. + depth | 5.74 | 53.0 | 44.0 | 39.0 | — |
| Ego²-Map | Pano. + odo. + depth | 5.54 | 56.0 | 47.0 | 41.0 | — |
| DreamWalker | Pano. + odo. + depth | 5.53 | 59.0 | 49.0 | 44.0 | — |
| AO-Planner | Pano. + depth | 5.55 | 59.0 | 47.0 | 33.0 | — |
| g3D-LF | RGB + odo. + depth | 5.70 | 59.5 | 47.2 | 34.6 | — |
| Seq2Seq | RGB + depth | 7.77 | 37.0 | 25.0 | 22.0 | — |
| NaVid-4D | RGB + depth | 5.99 | 55.7 | 43.8 | 37.1 | — |
| NavMorph | RGB + depth | 5.75 | 56.9 | 47.9 | 33.2 | — |
| NaVid | Single RGB | 5.47 | 49.1 | 37.4 | 35.9 | 953K |
| Sim2Real | Single RGB | 5.95 | 55.8 | 44.9 | 30.4 | 0K |
| StreamVLN* | Single RGB | 5.98 | 51.3 | 45.6 | 42.3 | — |
| Uni-NaVid | Single RGB | 5.58 | 53.3 | 47.0 | 42.7 | 3,577K |
| NaVILA* | Single RGB | 5.37 | 57.6 | 49.7 | 45.5 | 12,574K |
| JanusVLN* | Single RGB | 5.17 | 58.0 | 52.8 | 49.2 | 0K |
| AdaGeoVLN | Single RGB | 5.27 | 60.7 | 55.7 | 51.4 | 0K |
| Method | Observation | NE ↓ | SR ↑ | SPL ↑ | nDTW ↑ | External data |
|---|---|---|---|---|---|---|
| VLN⟳BERT | Pano. + odo. + depth | 8.98 | 27.0 | 22.6 | 46.7 | — |
| AO-Planner | Pano. + depth | 7.06 | 43.3 | 30.5 | 50.1 | — |
| Seq2Seq | RGB + depth | 12.10 | 13.9 | 11.9 | 30.8 | — |
| NavMorph | RGB + depth | 8.85 | 30.8 | 22.8 | 44.2 | — |
| Sim2Real | Single RGB | 8.79 | 36.7 | 25.5 | 18.1 | 0K |
| StreamVLN | Single RGB | 6.72 | 48.6 | 42.5 | 60.2 | — |
| Uni-NaVid | Single RGB | 6.24 | 48.7 | 40.9 | — | 3,577K |
| NaVILA | Single RGB | 6.77 | 49.3 | 44.0 | 58.8 | 13,132K |
| JanusVLN* | Single RGB | 6.46 | 51.4 | 44.3 | 59.1 | 0K |
| AdaGeoVLN | Single RGB | 5.64 | 54.1 | 44.7 | 61.8 | 0K |
NE is navigation error in meters; OS is oracle success; SR is success rate; SPL is success weighted by path length; nDTW is normalized dynamic time warping. All metrics except NE use a 0–100 scale. Pano.: panoramic observations; odo.: odometry. External data counts samples beyond R2R/RxR; — denotes an unreported total or metric. R2R-CE StreamVLN* is the oracle-navigation ablation; RxR-CE uses standard StreamVLN, whose total external sample count including EnvDrop is not separately quantified. NaVILA* excludes human-following data; JanusVLN* uses no additional external data.
SR gains over JanusVLN* on R2R-CE and RxR-CE, respectively. SPL improves by 2.2 and 0.4 points. On RxR-CE, nDTW increases by 2.7 points and NE decreases by 0.82 m. All use a single RGB stream, with no additional navigation training samples beyond R2R/RxR.
The repeated-deep control uses the same three fusion locations, but injects the terminal VGGT representation at every location.
| Geometry configuration | SR ↑ | SPL ↑ | OS ↑ | NE ↓ |
|---|---|---|---|---|
| No geometry · Qwen3.5 SFT | 46.5 | 42.3 | 56.0 | 5.19 |
| Single deep · G₂₃ → L₂ | 48.9 | 44.1 | 56.0 | 5.28 |
| Deep × 3 · G₂₃ | 42.1 | 36.9 | 56.7 | 6.64 |
| Hierarchical · G₁₁ / G₁₇ / G₂₃ | 55.7 | 51.4 | 60.7 | 5.27 |
Hierarchical fusion improves SR/SPL by 9.2/9.1 points over no geometry, 6.8/7.3 over single-deep fusion, and 13.6/14.5 over Deep × 3 at matched fusion locations. The geometry-free policy retains a slightly lower NE (5.19 versus 5.27 m).
04 / Memory efficiency
Selected retention strategies on R2R-CE Val-Unseen compare navigation accuracy with measured GFM-KV and allocated GPU memory. Token budgets and measured memory footprints are reported separately.
Mean retained VGGT global-attention KV memory on R2R-CE Val-Unseen.
Nav. 900K improves SR/SPL by 0.4/1.4 points over Hybrid Inc. (8+48), which achieves 55.3% SR / 50.0% SPL.
The 900K budget counts token slots across all 24 layers, rather than unique scene tokens. Each layer independently retains its highest-scoring candidates, with special camera/register tokens given retention priority.
Relative to Hybrid Inc. (8+48), mean allocated GPU memory decreases from 24,161.94 to 19,275.87 MB (20.22%); mean GFM-KV memory decreases from 5,767.55 to 3,475.92 MB (39.73%).
05 / Navigation in action
Four Habitat episodes, one from each scene, compare AdaGeoVLN against JanusVLN and Qwen3.5 SFT under the same navigation instruction.
VLN-CE visualization 1
Instruction: Take a right after the pool table and take your first left and walk into the room and wait in between the two rooms.
VLN-CE visualization 2
Instruction: Walk into the hall ahead. Walk down the flight of stairs with an iron railing. Continue to the bottom of the stairs and stop near the bench at the bottom next to a closed door.
VLN-CE visualization 3
Instruction: Exit the laundry room and walk up the stairs to the left. Continue past the carpeted stairs and turn right. Turn right again and enter the living room. Wait near the coffee table.
VLN-CE visualization 4
Instruction: Walk past the shelf and down the hallway to the left of the double doors. At the end of the hallway turn right and stop in front of the end table.
06 / Real-world deployment
A ZED X Mini supplies the RGB stream. An onboard NVIDIA Jetson AGX Orin handles image preprocessing and communication; policy inference runs remotely on an NVIDIA RTX A6000.