360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
Abstract
A new photorealistic urban benchmark reveals large performance gaps for embodied agents in city-scale navigation and spatial reasoning.
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.
Community
๐ #ECCV2026 ๐๐ง๐ญ๐ซ๐จ๐๐ฎ๐๐ข๐ง๐ ๐๐๐๐๐ข๐ญ๐ฒ๐๐ซ๐๐ง๐ ๐๐งญ
"Can AI truly understand and navigate a real city?"
We introduce 360CityArena, a realistic urban navigation benchmark built from 360ยฐ videos of Akihabara, Tokyo. ๐ฏ๐ต
๐๏ธ Our arena:
โ
602 real-world 360ยฐ videos*
โ
85 streets*
โ
175 navigation & spatial reasoning tasks
Get this paper in your agent:
hf papers read 2608.08814 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper