Three models in two months, and now a fourth. On July 24, Anthropic rolled out Claude Opus 5, a flagship built to fix the one complaint enterprise customers keep raising: token bills climbing faster than the actual value the models deliver.
Opus 5 costs exactly what Opus 4.8 did: $5 per million input tokens, $25 per million output tokens. Yet on Frontier-Bench and GDPval-AA, benchmarks built around real-world work tasks, its scores more than doubled. On ARC-AGI 3, which tests a model's ability to handle problems it hasn't seen before, Opus 5 scored three times higher than the next-best model. Running at maximum effort on CursorBench 3.2, it lands within 0.5% of Claude Fable 5's peak score, at half the price per task.
The headline feature is an effort dial: low, medium, high. Users decide whether to save tokens on a simple job or let the model grind longer on a hard one. There's also a Fast Mode, running roughly 2.5 times quicker at double the base rate. Developers can now swap available tools mid-conversation without blowing up the prompt cache, a small change that adds up to real savings at API scale.
There are caveats. Anthropic's own behavioral audits found Opus 5 has the lowest misalignment score of any model the company has shipped: it's less prone to gaming or misleading users. It still trails Claude Mythos 5 on cybersecurity exploitation tasks, though, and for genuinely autonomous, multi-step projects Anthropic itself points customers toward Fable 5 instead.
The timing isn't an accident. OpenAI made a similar pitch for cheap efficiency with GPT-5.6 earlier in July. The industry has hit a wall on charging ever more for raw intelligence: companies now pick models by what shows up on the monthly bill, not just by benchmark charts.



