Preprint reports a gap between multimodal agents and a human reference in 3D world auditing
WorldAuditBench is an arXiv preprint describing 213 anomaly tasks in 13 simulated interactive 3D environments. Its authors report that five tested multimodal models, assessed with two agent designs, achieved success rates of 6.6% to 42.3%, against a reported 83.4% human figure. The supplied record does not establish these findings as peer-reviewed or independently replicated. [1]
This edition passed Imananq's enhanced publication checks. Some material claims remain explicitly attributed to official or company sources because no independent source is currently bound to this edition. The engine continues checking approved sources and will add corroboration only through a new edition that passes the full gate.

An arXiv preprint introduces WorldAuditBench, a benchmark for testing whether multimodal AI agents can navigate simulated 3D environments and identify anomalies. Its authors report that the tested agent configurations scored below a reported human reference result, but important evaluation details are not included in the supplied record. [1]
01
What we know now
- 01
[1] arXiv, WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents, arXiv:2609.40325v1, submitted 30 September 2026. https://arxiv.org/abs/2609.40325v1
- 02
The supplied primary record is complete for the arXiv abstract page, but does not include task-level materials or a full evaluation protocol.
02
Anomaly tasks described by the authors
The preprint describes 213 tasks for auditing simulated interactive 3D worlds.Interactive 3D environments
The authors say the benchmark covers 13 environments.Reported agent success range
Across five evaluated models and two approaches, the authors report success from 6.6% to 42.3%.Reported human reference result
The abstract reports 83.4%, but the supplied evidence does not describe the human-evaluation protocol.All figures are reported in the preprint and are not independently verified by the supplied record. [1]
03
What the benchmark is intended to test
The authors present WorldAuditBench as a benchmark for 3D world auditing: locating and validating anomalies while moving through a simulated environment. In their framing, an agent must both choose actions to search the world and use visual reasoning to judge what it observes. [1]
That is a bounded test of multimodal agents in interactive simulated worlds, rather than a broad measure of AI capability. The supplied abstract does not name the five anomaly families or provide task-level materials. [1]
- The authors describe 213 anomaly tasks across 13 interactive 3D environments.
- The benchmark is said to span five anomaly families.
- Examples in the abstract include floating objects, traversable walls, and scene-inconsistent objects.
04
How the authors say they evaluated agents
The preprint reports two auditing approaches. The first separates exploration from anomaly identification. The second uses an end-to-end vision-language-model agent, where visual reasoning guides the next action. [1]
The authors say they tested five frontier models under a fixed exploration budget. The supplied record does not identify those models, describe their configurations, or state the budget, so it cannot support a detailed comparison of individual systems or approaches. [1]
- Two-stage approach: vision-language-action exploration, then vision-language-model anomaly identification.
- End-to-end approach: a vision-language-model agent uses visual reasoning to guide action selection.
- Five frontier models were evaluated under a fixed exploration budget, according to the authors.
05
Reported results and what remains unknown
Across the tested models and two approaches, the authors report success rates from 6.6% to 42.3%, compared with a reported human figure of 83.4%. They present the gap as evidence of current limitations in gathering and interpreting evidence during exploration. [1]
These are author-reported findings in a preprint. The supplied record does not establish them as peer-reviewed or independently replicated, and it lacks the task-level materials and evaluation details needed to assess how robust the comparison is. [1]
- Reported agent success ranged from 6.6% to 42.3%.
- Reported human performance was 83.4%.
- The supplied abstract does not explain the human-study method or which configurations produced each score.
06
Why the report is worth following
The work sets out a focused test for an agent-design problem: using visual observations to decide where to search, then checking a suspected anomaly. The authors' reported results suggest that the tested configurations did not reach their reported human reference result on this benchmark. [1]
Readers can consult the primary arXiv record for WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents. It records the preprint's title, authors, category, submission date, benchmark description, and reported results. [1]
- Primary source: arXiv:2609.40325v1.
- Recorded submission date: 30 September 2026.
- Category: cs.AI preprint.
07
What to check before relying on the results
The preprint is an early benchmark report. Its performance figures are author-reported, and the supplied record does not establish peer review or independent replication. [1]
- 01
Read the primary record for version 1: arXiv:2609.40325v1. [1]
- 02
Do not treat these scores as a general measure of real-world agent reliability: the reported tasks concern simulated interactive 3D environments. [1]
- 03
Look for later materials that define the anomaly families, model configurations, exploration budget, human-evaluation protocol, and any code, task data, or environment access. [1]
08
Limits of this edition
The paper is recorded by arXiv as a cs.AI preprint, submitted on 30 September 2026. The supplied record does not establish it as peer-reviewed or independently replicated. [1]
The supplied abstract does not name the five anomaly families, identify the five models, or fully specify the fixed exploration budget. [1]
The supplied evidence does not explain the protocol, confidence intervals, or configuration-level results behind the reported 83.4% human-performance figure. [1]
The supplied record does not provide a benchmark-use pathway, public code, task data, or environment access. [1]
SRC
Source desk
Direct links to the material behind this selection. Seeing the source matters as much as reading the synthesis.


