THE SIGNAL IN ONE SENTENCE
A chair floating above the floor is not a philosophical problem. It is a bug. So is a wall that looks solid until a tester walks through it. So is a road sign that exists when you drive past, disappears when you turn around and returns the moment nobody is looking directly at the thing. Interactive worlds can be beautiful, persuasive and thoroughly broken. WorldAuditBench turns that familiar quality-assurance headache into an agent test. The new preprint contains 213 tasks spread across 13 interactive 3D environments built with Unreal Engine 5 and Three.js. Each task plants one labeled anomaly inside a scene and asks an auditor to explore from a first-person view, identify the defect and submit visual evidence supporting the report. The benchmark matters because finding a defect is not the same as recognizing one in a screenshot. Some problems are visible immediately. A crate may hover, two objects may overlap or a chair may be comically large. Other failures need an experiment. A fence has to be approached and tested before anyone knows its collision is missing. An object has to be viewed from several angles before its appearance change becomes clear. A location has to be revisited before a disappearing sign can be proven missing. That creates a loop familiar to every good tester: notice, suspect, investigate, compare and report. The agent must decide where to go before it knows exactly what is wrong. It must spend a limited action budget gathering evidence. It must remember what the scene looked like earlier. Then it must distinguish a genuine defect from an unusual but valid design choice. This is less like naming objects in a picture and more like entering a strange hotel room, tapping the walls and keeping receipts. The authors divide the anomalies into five families. Static physics covers unsupported, intersecting or incorrectly scaled objects. Interactive physics covers broken obstruction and contact behavior. Spatial consistency covers objects that change across positions or viewpoints. Temporal consistency covers changes across time. Semantic consistency covers arrangements that do not fit an object's function or the historical setting. Those five families contain 15 categories. The task collection includes 126 Unreal Engine 5 tasks and 87 Three.js tasks across indoor, urban, historical, industrial and natural settings. The authors began with scenes without the target bug, defined the exploration boundaries and interactions, then inserted an anomaly. Independent reviewers inspected each task. A configuration was accepted only after at least two distinct reviewers rated that same version as a pass. The result is deliberately controlled. Every main evaluation task contains one target anomaly. The agent receives a scene description, an anomaly-type hint and a matching example, but not the target object's identity, location or evaluation rubric. That is useful for isolating behavior. It is also friendlier than a real production world, where several bugs can coexist, the category may be unknown and nobody puts a little card beside the build saying the ghost lives somewhere in temporal consistency. The paper compares two ways to audit. In the interactive vision-language model setup, one multimodal agent observes the scene, chooses actions, investigates suspicions and writes the report. It can move, turn, look, interact, wait, retrieve older frames, review its action history, take notes and revise a finding. It receives 40 environment actions. The second setup separates movement from analysis. A vision-language-action model explores for 60 simulated seconds, producing a recorded trajectory. A separate vision-language model then watches that trajectory and looks for the anomaly. This is faster and cheaper because the analyst is not steering the search. It is also stuck with whatever evidence the explorer happened to collect. If the camera passes a suspicious fence and never tests it, the analyst cannot yell at the past to turn around. Five frontier model backbones were evaluated in both arrangements, each through its associated agent harness. The paper reports success rates from 28.2 to 42.3 percent for three of the stronger interactive auditors, while the full interactive range ran from 8.5 to 42.3 percent. The two-stage auditors ranged from 6.6 to 17.4 percent. Human auditors reached 83.4 percent. That human number needs its own label. The human baseline involved about ten computer science doctoral students. Two participants evaluated each task and could explore for up to ten minutes. Their reports were scored against the same rubrics by the same model judge used for agent reports. This is not a random sample of game testers, players or ordinary workers. It is a controlled expert baseline inside the benchmark. The judge is another measurement dependency. The authors used a frontier vision-language model to compare each report and its cited frames with the hidden task rubric. On a validation subset of human reports, its binary decisions agreed with separate human judgments 91.72 percent of the time, with Cohen's kappa of 0.645. That supports the judge as a practical instrument. It does not make every decision unquestionable, especially when a report is partially right, evidence is ambiguous or the same model family appears elsewhere in the experiment. The category results show where the floor gives way. Static physics and semantic inconsistencies were easier because a single frame can sometimes reveal the problem. Temporal consistency was brutal. Every evaluated model and paradigm scored 12.5 percent or lower on that family. A before-and-after comparison demands memory, a useful revisit and enough discipline to connect two observations that may be separated by many actions. Agents can look directly at both halves of a fact and still fail to assemble the sentence. The interactive approach's advantage was not only better reasoning after exposure. It reached relevant evidence more often. In one analysis using the strongest interactive auditor, geometric coverage of anomaly locations was 91.1 percent, compared with 47.4 percent for the separate explorer. Human-rated exposure, which required more than merely pointing the camera near the right place, was 68.1 percent versus 34.7 percent. When both approaches did expose the anomaly, the interactive auditor still performed better. Steering the investigation and connecting evidence both mattered. This distinction is the paper's most useful idea. A system can fail because it never reached the defect, because it saw the defect without recognizing it or because it recognized something strange but could not support the claim. Those are different engineering problems. Better navigation helps the first. Better visual reasoning helps the second. Memory, controlled tests and evidence management help the third. One overall score turns all three into soup. More time was not a miracle cure. On 126 Unreal tasks with one model, cutting the interactive budget from 40 actions to 20 reduced success from 33.3 to 23.0 percent. Increasing the budget to 60 actions raised it only to 34.9 percent. Giving an agent more wandering room helps a little after a point, but a robot taking the scenic route through a haunted shopping mall is still taking the scenic route. The hints mattered much more. Removing the example while keeping the anomaly-type hint reduced interactive success from 33.3 to 31.0 percent. Removing both the example and type hint dropped it to 15.9 percent. Knowing what kind of evidence to seek sharply changes exploration. That result is practical and cautionary. Specialized auditors may become useful when a team can tell them to check collision, persistence or scale. It does not show that an agent can enter an unfamiliar world and discover the entire defect taxonomy by itself. Real builds also contain more than one problem. The authors created smaller multi-anomaly trials around 21 Unreal tasks. The tested interactive agent was told the total number of anomalies, but only received the anchor anomaly's type and example. It found every target in three of 42 two-anomaly runs and two of 42 three-anomaly runs. Target recall stayed below 30 percent in both settings. The agent could pursue the known bug while the rest of the world quietly caught fire behind it. For game development, the immediate application is not replacing quality-assurance teams. It is widening the sweep. An automated auditor could revisit a nightly build, test doors and barriers, compare object states across checkpoints and package suspicious cases with screenshots and exact action traces. Human testers could then spend less time reproducing vague reports and more time judging severity, player impact and the weird interactions nobody thought to encode. Simulation teams have a second reason to care. A navigation or robotics policy can appear successful because the environment is defective. If an agent reaches a destination by passing through a wall, the score measures an exploit as though it were competence. Auditing the world before evaluating the policy separates model ability from simulator leakage. That becomes more important as generated 3D scenes are used for training. A machine that manufactures synthetic worlds can also manufacture synthetic shortcuts. Digital twins, architectural walkthroughs, industrial training systems and autonomous-vehicle simulations face the same basic risk. A visual model may treat a polished render as evidence that the underlying behavior is sound. World auditing asks for the opposite habit. Touch the barrier. Revisit the switch. Compare the readings. Preserve the observation that supports the claim. A realistic-looking world should be treated like an untested appliance, not like a witness under oath. The benchmark should not be stretched into a claim about physical safety. Its tasks are hand-built, simulated and bounded. Movement and interaction occur through a defined tool interface. The worlds contain one known target anomaly in the main evaluation. A warehouse, hospital or public street contains people, partial observability, sensor noise, changing rules and consequences that cannot be reset with a browser refresh. Success here would be encouraging evidence about investigation, not a license to send an unsupervised robot into the loading dock. The model comparison also bundles models with harnesses. Different command-line agents manage context and tool calls for different backbones. The paper uses a shared interface, but a result still reflects the model, its harness and the experimental setup together. Treating the table as a pure ranking of model intelligence would be cleaner than the experiment actually permits. There is a release caveat too. The authors say they will release evaluation code, task configurations and runnable environment packages subject to third-party licenses. At publication, this remains a new preprint and an author-created benchmark. The public evidence is detailed, but independent teams still need to reproduce the results, inspect package availability and test whether the tasks predict performance on their own worlds. The best production design is a layered inspection process. Start with cheap deterministic checks for impossible transforms, missing collisions, broken navigation meshes, asset references and state transitions. Add scripted probes for known behaviors. Use an agent for open-ended exploration and evidence collection. Keep a human responsible for ambiguous reports, player experience and release decisions. Every finding should preserve the build version, starting state, action sequence, before-and-after frames, suspected category and a reproducible test. An agent report without evidence is a rumor wearing a lanyard. A useful auditor should show how it reached the location, what expectation it tested, which action exposed the failure and whether the result repeated. If the issue appears only after a specific viewpoint or delay, those conditions belong in the report. If the agent cannot reproduce the problem, uncertainty belongs there too. The plain signal is that multimodal agents can already perform pieces of interactive quality assurance, but they are not yet reliable world inspectors. The strongest reported result still missed most target anomalies. The larger gap is not eyesight alone. It is deciding where to look, testing the right hypothesis, remembering what changed and attaching enough evidence that another person can trust the report. In a simulated world, as in the ordinary one, noticing something weird is cheap. Proving what broke is the job.
01
WHAT ACTUALLY CHANGED
WorldAuditBench was submitted to arXiv on September 30, 2026.
The benchmark contains 213 manually constructed anomaly tasks across 13 interactive 3D environments and 27 bounded scenes.
Unreal Engine 5 supplies 126 tasks and Three.js supplies 87 tasks.
The tasks span indoor, urban, historical, industrial and natural settings.
Five anomaly families cover static physics, interactive physics, spatial consistency, temporal consistency and semantic consistency.
Those families are divided into 15 anomaly categories.
Each main evaluation task contains one planted, labeled anomaly.
Agents receive a scene description, anomaly-type hint and matching example, but not the target object, location or evaluation rubric.
Interactive auditors can navigate, inspect, wait, use memory and revise reports within a 40-action budget.
Two-stage auditors analyze a fixed trajectory collected by a vision-language-action model during 60 seconds of simulated exploration.
Across the evaluated systems, reported success ranged from 6.6 to 42.3 percent.
Human auditors reached 83.4 percent under a different time budget of up to ten minutes per task.
Temporal-consistency success was 12.5 percent or lower for every evaluated model and paradigm.
The strongest interactive setup exposed and identified more anomalies than the separated exploration-and-analysis pipeline.
In the smaller multi-anomaly tests, complete discovery occurred in only five of 84 runs across the two-anomaly and three-anomaly conditions.
02
WHY THIS MATTERS
A convincing 3D world can hide physical, spatial, temporal and semantic defects that static image inspection cannot expose.
Interactive testing requires the agent to decide what evidence to gather before it knows exactly where the defect is.
Failing to reach a bug, failing to recognize it and failing to support a report are separate problems that need separate fixes.
Adaptive investigation outperformed analysis of a prerecorded exploration because suspicions could guide the next action.
Memory matters when a defect appears only across viewpoints or before-and-after states.
Simulator bugs can make an embodied policy look capable when it is merely exploiting a broken environment.
Generated training worlds need auditing so synthetic data does not quietly reward impossible behavior.
A category hint sharply improved performance, suggesting narrow specialist auditors may arrive before general autonomous testers.
The multi-anomaly results show that finding a known target does not imply broad coverage of a messy build.
Human testers remain necessary for ambiguous behavior, severity, player experience and release judgment.
Evidence-rich reports can reduce reproduction work even when an agent is not trusted to close a bug.
Digital twins and safety simulations need world validation before their results can support decisions about physical systems.
03
WHERE IT COULD HELP
- Run nightly exploration sweeps against new game and simulation builds.
- Test barriers, doors, switches, object persistence and state transitions with scripted and agentic probes.
- Capture exact starting state, actions, frames and build identifiers for every suspected anomaly.
- Use anomaly-specific prompts for collision, scale, persistence, viewpoint and semantic checks.
- Revisit landmarks and compare saved observations when testing temporal and spatial consistency.
- Require a reproducible action sequence before promoting an agent finding to a confirmed defect.
- Route uncertain or low-evidence reports to human quality-assurance staff.
- Audit simulation environments before scoring navigation, robotics or embodied-agent policies inside them.
- Check generated 3D training scenes for shortcuts that reward impossible movement or interaction.
- Combine deterministic engine checks with open-ended visual exploration rather than choosing only one.
- Track exposure, recognition and evidence quality as separate evaluation metrics.
- Measure coverage across scene regions, interactions, viewpoints and time, not only bug counts.
- Use a second judge or human review for high-impact and ambiguous findings.
- Preserve failed searches so future agents can learn which regions and tests were already attempted.
- Test multiple simultaneous anomalies before claiming production-level coverage.
- Validate performance on the team's own worlds instead of assuming one benchmark transfers cleanly.
KEEP A HAND ON THE WHEEL
This article covers a September 30 author-produced preprint and benchmark. Main tasks contain one planted anomaly and supply an anomaly-type hint plus a matching example. Models use associated agent harnesses, so the comparison does not isolate model weights alone. Interactive and two-stage paradigms have different exploration mechanisms and budgets. The human baseline involved about ten computer science doctoral students with up to ten minutes per task. Reports were scored by a model judge that reached 91.72 percent agreement with separate human judgments on a validation subset, not perfect agreement. Multi-anomaly testing used a smaller Unreal-only slice. The authors state that code, configurations and runnable packages will be released subject to licenses; independent reproduction and transfer to real production worlds remain open questions. Simulated benchmark performance is not evidence of physical-world safety.
04
TERMS WORTH KEEPING
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on October 1, 2026.
PUBLICATION RECEIPT: Revision 1. Published October 1, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US