ModelsCulture OpenAI

GPT-6 Astra cleared all 48 levels of the I'm Not a Robot game

Illustration for the GPT-6 Astra CAPTCHA game story

The test built to prove a human is present has been completed end to end by a model, four days after that model shipped.

What was actually done

OpenAI’s Sharif Shameem posted a recording on September 7, 2026 showing GPT-6 Astra clearing all 48 levels of I’m Not a Robot, a browser game built by Neal Agarwal.

The game opens with the familiar checkbox and escalates from there: distorted text, image grids, reversed rules, puzzles, drawing tasks and interactive challenges that require reading context and intent rather than pattern matching. Astra worked through it with computer use, reading the screen and acting on it directly instead of answering questions about it.

The caveat, stated plainly

Agarwal’s game is a parody of CAPTCHA. It is not the anti-bot infrastructure that actually guards websites, and clearing all 48 levels does not mean those systems have been defeated.

The distinction matters less than it first appears. Production verification moved to behavioral and device signals years ago, precisely because models kept getting better at puzzles. What a run like this marks is not a break in web security but a public reading of how far computer use has moved, on a task nobody designed as a benchmark.

Why an unofficial task reads better than a benchmark

Agarwal built the game as a joke about a familiar annoyance, which is exactly what makes the result legible. Nobody optimized against it, no lab published a score on it, and the levels escalate in ways that have no clean category: a drawing task sits next to a logic puzzle sits next to a challenge that only resolves if the player understands what the challenge is for.

Working through that sequence requires reading the screen, deciding what the current level wants, acting, and checking the result, repeatedly, without a human resetting the state between attempts. Those are the same operations that make computer use commercially interesting, packaged in a form a general audience can evaluate at a glance.

The result is best read as a demonstration rather than a measurement. It says nothing precise about reliability across a working day, which is the property that determines whether these systems are useful. It says a good deal about which categories of task have stopped being reserved for people.

Sources

ANOTHER News is published by ANOTHER, an AI-native content agency. Daily coverage also runs on Instagram.