What changed
Claude Opus 5's SWE-bench Pro entry appeared July 25 (79.2%), completing the benchmark picture after the July 24 launch. Across four key developer benchmarks, Opus 5 leads the frontier on reasoning and agentic coding while matching Fable 5 on software engineering tasks โ at half the price.
TL;DR
Claude Opus 5 scores 43.3% on Frontier-Bench (agentic coding, leads all models), 30.2% on ARC-AGI-3 (>3ร the prior record of 7.8% set by GPT-5.6 Sol), 96% on SWE-bench Verified, and 79.2% on SWE-bench Pro (vs Fable 5 at 80.3%) โ at $5/$25 per MTok vs Fable 5's $10/$50.
Developer signal
Four benchmarks, one clear takeaway: Opus 5 is optimized for agency and reasoning, not general-purpose breadth. (1) Frontier-Bench v0.1 (74-task agentic coding benchmark, successor to Terminal-Bench 2.1): Opus 5 at 43.3% max effort leads GPT-5.6 Sol (37.5%) and Fable 5 (33.7%) by a 6โ10 point margin. If you are building agentic coding pipelines or computer-use workflows, Opus 5 at $5/$25 is now the primary option to benchmark โ it leads the frontier and costs half of Fable 5. (2) ARC-AGI-3 (novel reasoning; no memorization possible by design): Opus 5 at 30.2% (High effort) is more than three times the previous best of 7.8% by GPT-5.6 Sol running at Max effort. ARC Prize confirmed it solved five previously unbeaten environments. The caveat: ARC-AGI-3 was evaluated at High effort only, not Max, due to the short testing window โ the Max-effort score is not yet available. (3) SWE-bench Pro (harder version of SWE-bench Verified; leaderboard updated July 25): Opus 5 at 79.2% is third behind Mythos 5 (80.3%) and Fable 5 (80.0%), and ahead of its predecessor Opus 4.8 (69.2%) by a full 10 percentage points. The Fable 5 vs Opus 5 gap on SWE-bench Pro (1.1 pp) is narrower than on Frontier-Bench (9.6 pp in Opus 5's favor) โ suggesting different optimization targets. (4) SWE-bench Verified: Opus 5 at 96%, Mythos 5 at 95.5%, Fable 5 at 95% โ the three Anthropic models cluster at the top. Practical guidance: for long-running agentic coding agents and computer-use tasks, switch to Opus 5; for knowledge-work breadth, Fable 5 remains the default. The price differential ($5/$25 vs $10/$50) is now decisively in favor of Opus 5 for agentic use cases.