Same Model, 30% Or 100%. The Harness Decided.
NVIDIA wrapped Claude Opus 5 in its own agent and cleared every public ARC-AGI-3 level, on a benchmark where the model alone scores about 30%.

On August 21st NVIDIA reported that AVO, its research agent built on Agentic Variation Operators, scored 100.00 on the public set of ARC-AGI-3, the interactive benchmark where an agent enters game-like environments with no instructions, no stated rules, and no stated goal. It completed all 25 environments and all 183 levels in 6,624 environment actions. The model inside was Claude Opus 5. ARC Prize separately reports roughly 30% for the same model at High reasoning effort.
What changed was everything around the model. AVO carries persistent memory of prior attempts and results, and runs a supervisor that watches the whole trajectory and redirects the main agent when progress stalls. The same architecture first ran seven days unattended on attention kernels, explored more than 500 directions, and beat cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5% on B200s. For ARC it saw each frame as a 64-by-64 text grid, no images at all, and used about 12% fewer actions than VISTA running the same model. NVIDIA is careful on both counts: this is the public set, not the semi-private or private ones, and the cross-system comparison *"should not be interpreted as a controlled ablation."*
Last week Anthropic said its own task-based evaluations *"no longer capture increases in models' capabilities."* This week a hardware company showed the same gap from the outside. A benchmark score is a measurement of a system, and the system is mostly not the model. That cuts two ways. Every leaderboard quoting a bare model number is measuring the wrong object, and any safety regime that governs capability by deciding which weights leave the building is governing a part that, on this evidence, accounts for about a third of the result.










