| Title: |
FOM-Nav: Frontier-Object Maps for Object Goal Navigation |
| Authors: |
Chabal, Thomas; Chen, Shizhe; Ponce, Jean; Schmid, Cordelia |
| Contributors: |
Models of visual object recognition and scene understanding (WILLOW); Département d'informatique - ENS-PSL (DI-ENS); École normale supérieure - Paris (ENS-PSL); Université Paris Sciences et Lettres (PSL)-Université Paris Sciences et Lettres (PSL)-Institut National de Recherche en Informatique et en Automatique (Inria)-Centre National de la Recherche Scientifique (CNRS)-École normale supérieure - Paris (ENS-PSL); Université Paris Sciences et Lettres (PSL)-Université Paris Sciences et Lettres (PSL)-Institut National de Recherche en Informatique et en Automatique (Inria)-Centre National de la Recherche Scientifique (CNRS)-Centre Inria de Paris; Institut National de Recherche en Informatique et en Automatique (Inria); Université Paris Sciences et Lettres (PSL)-Université Paris Sciences et Lettres (PSL)-Institut National de Recherche en Informatique et en Automatique (Inria)-Centre National de la Recherche Scientifique (CNRS); Center for Data Science NYU (CDS); New York University New York (NYU); NYU System (NYU)-NYU System (NYU); Courant Institute of Mathematical Sciences New York (CIMS); This work was granted access to the HPC resources of IDRIS under the allocation 2021-AD011012725 made by GENCI. It was supported in part by the French government under management of Agence Nationale de la Recherche as part of the “France 2030” program, PR AI RIE-PSAI projet, reference ANR-23-IACL-0008. JP was supported in part by the Louis Vuitton/ENS chair in artificial intelligence and a Global Distinguished Professorship at the Courant Institute of Mathematical Sciences and the Center for Data Science at New York University.; ANR-23-IACL-0008,PR AI RIE-PSAI,PR AI RIE-PSAI - Paris School of Artificial Intelligence(2023) |
| Source: |
https://hal.science/hal-05392088 ; 2025. |
| Publisher Information: |
CCSD |
| Publication Year: |
2025 |
| Subject Terms: |
Robotics; Computer vision; Robot navigation; Vision-language models; Semantic mapping; Scene exploration; [INFO.INFO-RB]Computer Science [cs]/Robotics [cs.RO]; [INFO.INFO-CV]Computer Science [cs]/Computer Vision and Pattern Recognition [cs.CV] |
| Description: |
This paper addresses the Object Goal Navigation problem, where a robot must efficiently find a target object in an unknown environment. Existing implicit memory-based methods struggle with long-term memory retention and planning, while explicit map-based approaches lack rich semantic information. To address these challenges, we propose FOM-Nav, a modular framework that enhances exploration efficiency through Frontier-Object Maps and vision-language models. Our Frontier-Object Maps are built online and jointly encode spatial frontiers and fine-grained object information. Using this representation, a vision-language model performs multimodal scene understanding and high-level goal prediction, which is executed by a low-level planner for efficient trajectory generation. To train FOM-Nav, we automatically construct large-scale navigation datasets from real-world scanned environments. Extensive experiments validate the effectiveness of our model design and constructed dataset. FOM-Nav achieves state-ofthe-art performance on the MP3D and HM3D benchmarks, particularly in navigation efficiency metric SPL, and yields promising results on a real robot. |
| Document Type: |
report |
| Language: |
English |
| Availability: |
https://hal.science/hal-05392088; https://hal.science/hal-05392088v1/document; https://hal.science/hal-05392088v1/file/FOMNav.pdf |
| Rights: |
http://creativecommons.org/licenses/by/ ; info:eu-repo/semantics/OpenAccess |
| Accession Number: |
edsbas.EE35AEC1 |
| Database: |
BASE |