Anthropic released Claude Opus 4.7 today, describing it as "a notable improvement on Opus 4.6 in advanced software engineering, with particular gains on the most difficult tasks." Pricing is unchanged from Opus 4.6: $5 per million input tokens, $25 per million output. It's available across all Claude products, the API, Amazon Bedrock, Google Vertex AI, and Microsoft Foundry.

The launch numbers are the vendor's, so treat them as directional. Anthropic's headline claim: "On our 93-task coding benchmark, Claude Opus 4.7 lifted resolution by 13% over Opus 4.6, including four tasks neither Opus 4.6 nor Sonnet 4.6 could solve." The company also says the model "demonstrates strong substantive accuracy on BigLaw Bench for Harvey, scoring 90.9% at high effort."

One timing note: this is the first model Anthropic has shipped since Project Glasswing, the critical-software security initiative it joined with AWS, Apple, Cisco, and others on April 7 — and the announcement puts noticeable weight on cybersecurity safeguards.

Point releases are where budgets change

A version number like 4.7 reads like a footnote. Nobody holds a meeting about it. But point releases are where AI budgets actually move, and this one shows both directions at once — at a price that didn't change.

The ceiling story is those four tasks. Every release quietly moves a few pieces of work from "impossible" to "possible," and no vendor announces which of your pieces just crossed the line. If your team tried something hard in the past six months — a messy legacy migration, a codebase nobody wanted to touch — the model that failed at it is no longer the one on sale. The 13% is a vendor stat; the four previously-unsolvable tasks are the shape of the thing that matters.

The floor story comes from customers, not benchmarks. Replit: "For the work our users do every day, we observed it achieving the same quality at lower cost—more efficient and precise at tasks like analyzing logs and traces." List price didn't move; per-task cost did, because the model wastes fewer tokens getting to the same answer. Quantium was blunter, calling it "the most capable model we've tested."

Neither story is a new paradigm, and that's the point. Roughly every six weeks, the same work gets cheaper and more reliable. Compounded, that cadence has one uncomfortable implication for anyone who ran an AI evaluation last year: the conclusion has a shelf life, and it has probably expired. We saw the same dynamic in February, when Opus 4.6 made agent teams practical — capability that had been demo-grade quietly became production-grade. Matching this quarter's models to this quarter's task list is most of the job, and it's why we scope engagements around the task, not the model.

What to do with this

  • Re-test one failed pilot. If AI stumbled on hard technical work in the past six months, re-run the exact same test on the current model before letting the old conclusion stand. Four previously-unsolvable tasks is precisely the kind of change old evaluations can't see.
  • Ask about per-task cost, not list price. Pricing held at $5/$25 per million tokens, but Replit's experience says the same work now takes fewer tokens to finish. If your AI bill is flat while usage grows, this effect is why.
  • Put an expiry date on AI conclusions. Any "we tried it, it can't" older than a quarter goes back into the queue — through the same three-question filter as every other idea on the list.

If a parked idea just crossed the line and you want a second opinion before funding it, that's a 30-minute conversation.

Source: Introducing Claude Opus 4.7 — Anthropic, April 16, 2026