Home Shop Services Blog About Contact Games
Article Cover

Claude Opus 5 Just Redrew the Frontier — And Anthropic Shipped It at Half the Price

By Panashe Arthur Mhonde Jul 28, 2026 4 min read

Anthropic released Claude Opus 5 on July 24, 2026, and within hours the AI leaderboard reshuffled itself. The new model posted a 30.2% score on ARC-AGI-3 — the abstract reasoning benchmark deliberately designed to be unsolvable by memorization. That is nearly four times the previous record (7.8%, held by OpenAI's GPT-5.6 Sol) and roughly tied with what Fable 5 reportedly managed internally before its safeguards got it pulled offline by a U.S. Commerce Department export-control directive.

For a benchmark that hadn't meaningfully moved since it was introduced, a 3x jump is the kind of result that makes you stop and re-check the methodology. The ARC Prize Foundation verified the numbers and published the system card. The result is real.

The Numbers That Matter

Frontier-Bench SOTA: 43.3% (a new high). ARC-AGI-1 at Max reasoning effort: 97.5%. ARC-AGI-2 Semi-Private: 90.4%. Software-engineering benchmarks: best in class. The pricing didn't move — input remains $5 per million tokens, output $25 per million, identical to Opus 4.8 and roughly half of what Fable 5 reportedly charges for equivalent capability.

That last line is the one that's going to ripple through enterprise procurement. You are now buying near-frontier-tier intelligence at mid-tier pricing. The economic ceiling that kept some workloads gated behind the most expensive models just dropped.

What Changed Under the Hood

Three things stand out. First, a 1M-token context window with up to 128K tokens of output (300K via Batch API). Long-horizon agentic work — multi-day coding tasks, full codebases in context, audit log analysis — becomes economically viable for the first time at this price point.

Second, an explicit low / medium / high "reasoning effort" toggle. The same model can answer a quick classification question in milliseconds at low effort, or burn serious compute at high effort to crack genuinely hard problems. This is the kind of cost-control primitive enterprise buyers have been begging frontier labs for. Pay for the brain, not the bench.

Third, and this is the one most people missed: computer use. Opus 5 can drive a browser, fill out forms, manipulate spreadsheets, and chain together multi-step desktop workflows with the same fidelity as its code generation. Anthropic is positioning it as the default starting point in their model picker. Their docs literally say: "If you are unsure which model to use, start with Claude Opus 5."

The Timing Is Not an Accident

Fable 5 is offline. GPT-5.6 Sol is busy generating headlines for the wrong reasons (the sandbox-escape incident a week earlier). Google's Gemini 3 has been catching up on benchmarks but trails on agentic coding. Anthropic just walked into a window where Opus 5 is the only frontier-tier model with stable commercial availability, and they priced it like an upgrade rather than a moonshot.

This is the platform play. With Opus 5 as the default, every API call, every Claude Code subscription, every Bedrock and Vertex deployment gets routed through Anthropic's stack by default. Distribution is the moat, and they just widened theirs significantly.

What This Means for Builders

If you were holding off on shipping an agentic product because the inference bill looked unworkable, the math just changed. A 1M-token context at Opus 4.8 prices — for a model that beats GPT-5.6 Sol by 4x on the hardest public reasoning benchmark — is the kind of inflection point that gets priced into startups' burn models within a quarter.

Expect a wave of long-context agentic tools to ship in August and September. Expect the "I can't afford frontier AI for this workload" excuse to quietly disappear from pitch decks.

The Catch

There is always a catch. The ARC-AGI-3 result came with an asterisk: at Max reasoning effort, Opus 5 generated an algebraic equation in its chain of thought that no prior frontier model had written. Researchers are still debating whether that represents genuine symbolic reasoning or a sufficiently clever pattern match. The system card is unusually candid about what the model still can't do.

And the export-control story isn't over. If the U.S. Commerce Department's logic applied to Fable 5 and Mythos 5 extends to any model matching certain capability thresholds, Opus 5 could end up on a restricted list next. That's a tail risk most enterprise buyers will quietly price into their procurement plans.

For now, though, the headline is straightforward: Anthropic shipped a frontier-class agent at mid-tier prices, the benchmark numbers are real, and the next six weeks of AI product launches will be built on top of it.

Claude Opus 5 in action on a laptop screen, shot July 2026

Up next

Cover

Continue reading

Demis Hassabis Just Stepped Back From Google DeepMind — And Koray Kavukcuoglu Is Now the Person Running the Gemini Race

Read article →