Anthropic released Claude Opus 4.6 today. The company's framing: "We're upgrading our smartest model. Across agentic coding, computer use, tool use, search, and finance, Opus 4.6 is an industry-leading model, often by wide margin." Pricing holds at $5 per million input tokens and $25 per million output, and the model is available on claude.ai, the API, and all major cloud platforms.
Two launch details matter more than the scoreboard. First, "Opus 4.6 features a 1M token context window in beta" — the context window being how much text the model can consider at once. Prompts over 200,000 tokens bill at a premium of $10 and $37.50 per million (Claude Platform only). Second, the model launched alongside agent teams in Claude Code — several AI agents coordinating on one job, the way a small team splits work — plus adaptive thinking, which lets the model decide how hard to think about each task.
On the scoreboard itself: Anthropic says that "On GDPval-AA, Opus 4.6 outperforms the industry's next-best model (OpenAI's GPT-5.2) by around 144 Elo points," and that "Opus 4.6 achieved the highest BigLaw Bench score of any Claude model at 90.2%." Those are day-one vendor numbers — directionally useful, independently unverified.
A dispatcher, not a chatbot
The line worth reading twice is smaller than the benchmarks. Anthropic reports that "Opus 4.6 autonomously closed 13 issues and assigned 12 issues to the right team members in a single day."
Parse that. The model didn't answer 25 questions — it worked a queue. It finished the 13 tasks it could finish and routed the other 12 to the right people. Closing tickets is automation; assigning them correctly is judgment about who does what. That second half is the job description of a dispatcher, and it's a different category of usefulness than any chatbot.
The 1M context window is what makes the dispatcher plausible. A million tokens is roughly 750,000 words — several years of customer emails, every quote you've sent, your full job history, in one sitting. Until now, pointing AI at your records meant feeding it snippets and hoping you picked the right ones. With the whole archive in view, "which of last year's 400 quotes did we underbid, and by how much?" becomes one question instead of a week of spreadsheet archaeology.
Agent teams push the same direction. Instead of one assistant doing one thing at a time, several coordinate: one reads the backlog, one drafts replies, one checks the drafts against your records. For a 2-to-50-person business, full-archive memory plus work-splitting is the shape of a back-office coordinator — the queue triaged, the routine handled, the exceptions routed to a named human.
To be clear about the altitude: that system is built deliberately, not switched on. On our Solution Levels scale, a dispatcher that touches real customer work sits near the top, and it earns that spot only after smaller wins have proven your data and your processes can support it.
What to do with this
- Ask one whole-archive question. Export a year of quotes, jobs, or support emails and ask something you'd never assign to a person — pricing drift, repeat complaints, which customers went quiet. One file, one afternoon, a real answer.
- Mind the meter on big prompts. The over-200k premium means full-archive runs cost real money each time. Treat them as weekly or monthly reviews, not per-task calls — and keep routine work on the cheaper Sonnet tier.
- Don't build the dispatcher first. "AI that runs my whole back office" is exactly the kind of idea that needs the three-question filter before anyone scopes it. Start with triage-plus-human-review on one queue, measure the routing accuracy, and promote it only when the numbers hold.
If you're wondering what a model with your whole archive in its head could do for your operation, the first conversation is 30 minutes.