MLLMs fail to sustain goal-directed navigation in real-scale 3D city

UrbanGround sandbox tests local perception and spatial reasoning in a replica of Hong Kong, finding that atomic abilities do not compose into reliable exploration.

PaperTop Universitycs.CVarXiv:2608.27456v1
Tianjie Ju · Zheng Wu · Yueqing Sun · Yuhan Cui · Bobo Li · Shengqiong Wu · +12 more

Shanghai Jiao Tong University · National University of Singapore · Meituan · The Chinese University of Hong Kong · Shanghai University

Research Digest··2 min read
The authors introduce UrbanGround, a sandbox built from 3D geospatial data of Hong Kong, to test whether multimodal large language models (MLLMs) can turn local visual perception into sustained navigation. They find that while agents show useful abilities in short-range spatial reasoning and scene recognition, they fail at orientation, pedestrian-aware movement, and correcting accumulated errors over extended exploration.

What they did

The authors constructed UrbanGround, the first sandbox to evaluate MLLM agents in a physically constrained replica of an entire real city (Hong Kong) built from territory-wide 3D geospatial data. The sandbox supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city, explore, and act based on local visual input. The analysis was structured around three research questions: whether agents can ground a local scene after active observation, whether that grounding supports navigation as destinations become farther and less explicit, and whether behavior survives changes in route availability and pedestrian motion.

Key findings

  • Agents show useful atomic abilities in visual recognition and short-range spatial reasoning, but orientation and pedestrian-aware movement remain unreliable.
  • The central failure mode is that local abilities do not compose into sustained goal-directed behavior over extended exploration; errors accumulate without effective correction.
  • Performance degrades significantly as destinations become more distant and less explicit, and when route availability or pedestrian motion changes.

Why it matters

The work highlights a critical gap between current MLLM capabilities and the demands of real-world urban navigation, where sustained reasoning, error correction, and adaptation to dynamic obstacles are essential. UrbanGround provides a reproducible testbed to probe these limitations, potentially guiding the development of more robust spatial agency in embodied AI systems.

Caveats

UrbanGround is a simulated environment; real-world urban navigation involves additional sensory modalities (e.g., audio, haptics) and physical constraints not captured here. The evaluation focuses on a single city (Hong Kong), and generality to other urban layouts is untested. The study does not propose new methods or training techniques, only diagnostic tasks.

§

Analysis

The paper aligns with a growing body of work that evaluates MLLMs in closed-loop, embodied settings, revealing limitations that are masked in static image benchmarks. By building a physically realistic city-scale environment, UrbanGround moves beyond toy simulators and highlights a compositionality problem: atomic competencies do not automatically yield reliable long-horizon behavior. This echoes challenges seen in other domains like robot manipulation and web navigation, where error accumulation and lack of automated correction are common failure modes.

newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.

MLLMs fail to sustain goal-directed navigation in real-scale 3D city | Zotpaper