Skip to content
D-CSIL

AI News · 2026-10-07 · 4:00 PM CT

AI learns to see: MIT's visual memory aces ARC-AGI-3

TL;DR

On October 1, an MIT team led by Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, and Kaiming He published VISTA, a visual harness that gives AI agents a lossless visual memory — every past frame stored in original form, retrievable mid-reasoning. With no model retraining and no game-specific code, Claude Opus 5.0 went from a 40.68 Relative Human Action Efficiency score to a perfect 100.00 across all 25 public ARC-AGI-3 games, using 57.4% fewer actions than first-time human players. GPT-5.6 Sol went from 13.33 to 99.00 under the same harness.

An astronaut figure playing a video game in a neon-lit dark room, monitor glowing
Photo: Erik Mclean / Pexels

The problem it solves

ARC-AGI-3 is a benchmark from the ARC Prize: 25 interactive visual games where the agent gets no rules and no goal description. It has to discover how each world works by playing. Most leading approaches fed the model a text version of the world — a 64x64 grid of numbers — and had it write programs (sometimes thousands of lines of Python) to build a simulator for trial and error.

The VISTA team argues the models were never bad at reasoning; they were blind. Vision-language models encode each image once into a compressed representation, and details they might need later get discarded or summarized away. VISTA reframes the games as a vision problem: the agent sees raw screenshots, and every frame the environment returns is archived losslessly so the model can flip back, zoom in, and re-read exact pixel values while it reasons.

The numbers

Claude Opus 5.0 with VISTA completed all 25 public games with a perfect RHAE score of 100.00, in 7,302 total actions — 57.4% fewer than the 17,135 actions first-time human players needed. GPT-5.6 Sol scored 99.00. The same models with the official text-based interface scored 40.68 and 13.33 respectively. The authors state that VISTA is, to their knowledge, the first vision-based system to reach perfect or near-perfect performance on ARC-AGI-3 without program synthesis.

The setup is deliberately minimal: a four-sentence prompt shared across every game, with nothing teaching the model how to play any specific one. There is no trained component in the harness itself — the intelligence comes from the model, and VISTA just gives it eyes and a memory.

Beyond the benchmark games

The same harness, with minimal adaptation, was applied to 34 browser games from GameWorld, 10 browser games from AI GameStore, and 39 visual tracking problems from BabyVision. On GameWorld, GPT-5.6 Sol with VISTA beat the novice human baseline; on AI GameStore it beat the human median. The pattern is consistent: when the bottleneck is perception, better perception beats bigger reasoning.

One caveat to keep straight: all the headline scores are on the 25 public games. The ARC Prize keeps private game sets specifically to check for overfitting, and the authors note their models postdate the public games, so generalization to unseen games is the next test.

Why this matters for builders

The lesson transfers directly to computer-use agents: stop assuming the model lacks intelligence when it struggles, and check what it is missing from its eyes and memory. If you are building an agent that clicks through a UI, the VISTA results suggest two cheap experiments before you fine-tune anything — give the agent raw screenshots instead of a compressed accessibility tree, and keep past frames retrievable at full resolution instead of summarizing them away.

The code is public on GitHub at github.com/joshhhhhan/VISTA. If you run browser or GUI agents, the four-sentence harness is a weekend's worth of prototyping to test whether your agent was blind.