The authors built SuperNav around a pretrained MLLM that is not fine-tuned for navigation.
Agent harness lets a pretrained multimodal LLM navigate any task without fine-tuning
SuperNav keeps the MLLM's general abilities intact and delegates motion execution to navigation tools, outperforming four baselines on instance-level, multi-object, and demand-driven benchmarks.
Top University
Jinkai Zhang · Jingyi Xu · Yuanhong Yu · Jiarui Guo · Ruizhen Hu · Hujun Bao · +2 more
Zhejiang University · Shenzhen University · Causa Robotics
Research Digest··2 min read
The authors present SuperNav, a navigation framework that equips a frozen pretrained multimodal large language model with an agent harness for tool use, progress tracking, and context management.
Why this paper
From Zhejiang University and 2 others
In one line
A pretrained multimodal LLM delegates decision-making to navigation tools to handle any navigation task in any scene.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§