Quantitatively evaluates O3D-SIM using the Matterport3D dataset and Success Rate metric in the Habitat simulatorQuantitatively evaluates O3D-SIM using the Matterport3D dataset and Success Rate metric in the Habitat simulator

Quantitative Evaluation of O3D-SIM: Success Rate on Matterport3D VLN Tasks

2025/12/16 11:01
4 min di lettura
Per feedback o dubbi su questo contenuto, contattateci all'indirizzo crypto.news@mexc.com.

Abstract and 1 Introduction

  1. Related Works

    2.1. Vision-and-Language Navigation

    2.2. Semantic Scene Understanding and Instance Segmentation

    2.3. 3D Scene Reconstruction

  2. Methodology

    3.1. Data Collection

    3.2. Open-set Semantic Information from Images

    3.3. Creating the Open-set 3D Representation

    3.4. Language-Guided Navigation

  3. Experiments

    4.1. Quantitative Evaluation

    4.2. Qualitative Results

  4. Conclusion and Future Work, Disclosure statement, and References

4.1. Quantitative Evaluation

To facilitate the construction of O3D-SIM and its quantitative evaluation, we employ the Matterport3D dataset [36] within the Habitat simulator [37]. Matterport3D, a comprehensive RGB-D dataset, encompasses 10,800 panoramic views derived from 194,400 RGB-D images across 90 large-scale buildings. It offers surface reconstructions, camera poses, and 2D and 3D semantic segmentations — critical components for creating accurate Ground-truth models. Both Matterport3D and Habitat are widely utilized for assessing the navigational abilities of VLN agents in indoor settings, enabling robots to execute navigational tasks dictated by natural language commands in a seamless environment, with performance meticulously documented. To evaluate O3D-SIM, we compiled 5,267 RGB-D frames and their respective pose data from five distinct scenes, applying this dataset across all mapping pipelines included in our assessment. Additionally, we gathered real-world environment data for evaluation purposes, thereby expanding our analysis to encompass six unique scenes.

\ Baseline: We evaluate the performance of the O3D-SIM against the logical baseline used in our previous work, VLMaps with Connected Components, and also evaluate against our approach SI Maps from [1]. The three methods mentioned for comparison are chosen as they try to achieve things similar to our approach.

\ Evaluation Metrics: Like prior approaches [2, 38, 39] in VLN literature, we use the gold standard Success Rate metric, also known as Task Completion metric to measure the success ratio for the navigation task. We choose Success Rate on navigation tasks as they directly quantify the overall approach and indirectly quantify the performance of O3D-SIM in detecting the instances because if the instances along the way are not properly detected, the queries are bound to fail. We compute the Success Rate metric through human and automatic evaluations. For automatic evaluation, we define success if the agent reaches within a threshold distance of the ground truth goal. Here, the agent’s orientation concerning the goal(s) doesn’t matter and might show success even when the agent fails. For example, if there are multiple paintings and the

\ Table 1. O3D-SIM outperforms the baseline methods from [1] by significantly large margins on the Success Rate metric. It also shows an improvement from our previous approach, i.e., SI-Maps. The best results are highlighted in bold. For this evaluation, the agent executes a set of open-set and closed-set queries in 5 different scenes from Matterport3D and 1 real word environment.

\ agent is asked to point to a particular painting at the end of a query, the agent may reach with a close distance of the desired painting but end up looking at something undesired. Hence, we also use human evaluation to verify if the agent ends up in a desired position according to the query. Human Verification takes in votes from the three people and decides, based on these votes, the success of a task.

\ Results: We present the results of the evaluation metric Success Rate in Table 1. In our experimentation, we observe a remarkable improvement in performance compared to the other approaches we have shown in our paper, especially against the baselines from [1]. O3D-SIM performs better than VLMaps with CC due to its ability to identify instances robustly. SI Maps and O3D-SIM perform better than the baselines due to their ability to separate instances. However, O3D-SIM has the edge over SI Maps due to its open set and 3D nature, allowing it to understand the surroundings better.

\

:::info Authors:

(1) Laksh Nanwani, International Institute of Information Technology, Hyderabad, India; this author contributed equally to this work;

(2) Kumaraditya Gupta, International Institute of Information Technology, Hyderabad, India;

(3) Aditya Mathur, International Institute of Information Technology, Hyderabad, India; this author contributed equally to this work;

(4) Swayam Agrawal, International Institute of Information Technology, Hyderabad, India;

(5) A.H. Abdul Hafez, Hasan Kalyoncu University, Sahinbey, Gaziantep, Turkey;

(6) K. Madhava Krishna, International Institute of Information Technology, Hyderabad, India.

:::


:::info This paper is available on arxiv under CC by-SA 4.0 Deed (Attribution-Sharealike 4.0 International) license.

:::

\

Disclaimer: gli articoli ripubblicati su questo sito provengono da piattaforme pubbliche e sono forniti esclusivamente a scopo informativo. Non riflettono necessariamente le opinioni di MEXC. Tutti i diritti rimangono agli autori originali. Se ritieni che un contenuto violi i diritti di terze parti, contatta crypto.news@mexc.com per la rimozione. MEXC non fornisce alcuna garanzia in merito all'accuratezza, completezza o tempestività del contenuto e non è responsabile per eventuali azioni intraprese sulla base delle informazioni fornite. Il contenuto non costituisce consulenza finanziaria, legale o professionale di altro tipo, né deve essere considerato una raccomandazione o un'approvazione da parte di MEXC.

Potrebbe anche piacerti

Claude Code has been found to have two caching bugs that could silently increase API costs by 10-20 times.

Claude Code has been found to have two caching bugs that could silently increase API costs by 10-20 times.

PANews reported on March 31 that, according to 1M AI News, a developer reverse-engineered a 228MB binary file of the standalone Claude Code installer using Ghidra
Condividi
PANews2026/03/31 11:37
US President Trump willing to end Iran war without reopening Strait of Hormuz – WSJ

US President Trump willing to end Iran war without reopening Strait of Hormuz – WSJ

The post US President Trump willing to end Iran war without reopening Strait of Hormuz – WSJ appeared on BitcoinEthereumNews.com. Citing administration officials
Condividi
BitcoinEthereumNews2026/03/31 11:02
Investors flock to IOTA miners in pursuit of stable returns

Investors flock to IOTA miners in pursuit of stable returns

The post Investors flock to IOTA miners in pursuit of stable returns appeared on BitcoinEthereumNews.com. After securing a preliminary victory in its protracted legal battle with the U.S. Securities and Exchange Commission (SEC), XRP (Ripple) has once again become a market focus. Within hours of the announcement, on-chain data revealed a discreet transfer of 15,000,000 XRP. While this amount is not significant compared to whale-level holdings, its timing and context have nonetheless drawn market attention: some analysts believe it may be related to liquidity reallocation, adjustments to cross-border payment channels, or early institutional investment. At the same time, market attention is gradually shifting from short-term price fluctuations to more sustainable profit models. Following the XRP legal victory, a large number of small and medium-sized investors have chosen the IOTA Miner cloud mining platform as an alternative to hedge against volatility and achieve stable returns. The platform’s core advantages include: Stable returns: Users receive a fixed daily mining reward regardless of market fluctuations; Low barriers to entry: No expensive hardware required; easy mobile participation; Risk hedging: Withdrawals are possible during price declines, effectively preventing significant losses; Environmentally friendly: The mining pool’s electricity is entirely sourced from renewable energy, making it efficient and sustainable. What is IOTAMiner? Founded in 2018 and headquartered in the UK, IOTAMiner is a reputable global cloud mining platform with seven years of experience, serving over 9 million users in over 100 countries. As the world’s first cloud mining platform integrating artificial intelligence with renewable energy, IOTAMiner maintains a strategic reserve of over 8,000 Bitcoins, operates in full compliance, and is committed to providing users with a 100% return on investment guarantee. IOTA Miner Registration Steps 1. Quick Registration Sign up in just a minute and receive a $15 newbie bonus to start earning immediately. 2. Link Your Wallet and Select Your Currency Link your wallet and select a major cryptocurrency (such as…
Condividi
BitcoinEthereumNews2025/09/18 02:02