OpenAI released GPT-5.5 and GPT-5.5 Pro today, calling the base model "our smartest model yet—faster, more capable, and built for complex tasks like coding, research, and data analysis across tools." It's rolling out to ChatGPT Plus, Pro, Business, and Enterprise plans and to Codex, with API access promised "soon."
The launch numbers, all per OpenAI: "GPT-5.5 scores 84.9% on GDPval, which tests agents' abilities to produce well-specified knowledge work across 44 occupations." It "reaches 78.7% on OSWorld-Verified, which measures whether a model can operate real computer environments on its own." And "On Terminal-Bench 2.0, it achieves a state-of-the-art accuracy of 82.7%." OpenAI also says the model "matches GPT-5.4 per-token latency in real-world serving, while performing at a much higher level of intelligence," and that it was "released with our strongest set of safeguards to date."
API pricing lands at $5 per million input tokens and $30 per million output. GPT-5.5 Pro, the heavier tier, costs $30 and $180 — six times the base rate. OpenAI's cost framing: "GPT-5.5 delivers state-of-the-art intelligence at half the cost of competitive frontier coding models."
From coding puzzles to a Tuesday's work
Most launch benchmarks measure things no business owner does. Competition math, terminal tasks, coding leaderboards — useful for ranking labs, useless for deciding whether a model can draft your proposals. GDPval is different. It's built from real occupational deliverables — the kind of output people in 44 different jobs actually produce — and OpenAI now puts it first in the pitch.
That 84.9% is the number to notice, and not because of its size. It marks a shift in what the industry measures. The old question was "can it solve a puzzle." The new one is closer to "can it do a Tuesday's worth of your work." When the yardstick becomes occupational output, a model release stops being a spectator sport for a 12-person plumbing company or a 30-person accounting firm. It becomes a price list for deliverables you currently produce by hand — at $5 per million input tokens.
But read OpenAI's sentence again and notice the quiet qualifier: well-specified. The model scored 84.9% on knowledge work that came with clear inputs and a clear definition of done. That qualifier is where automation projects live or die. In our experience, the bottleneck inside a small business is rarely model intelligence — it's that nobody has written down what "done" looks like for the task. Most AI ideas fail as use cases for exactly that reason, and no benchmark score fixes it. The specification work is yours either way.
The second thing worth noticing is the Pro tier. Same generation, six times the price: $30 and $180 against $5 and $30. That's OpenAI pricing task difficulty directly — pay a premium for the hardest slice of your work, standard rates for everything else. A menu like that rewards matching: route the routine 95% to the base tier and escalate only the cases that genuinely need more. It's the same discipline we apply through Solution Levels — not every task deserves the same tool, and the difference shows up directly in your costs.
One more thread: the 78.7% OSWorld-Verified claim continues the computer-use story we covered when GPT-5.4 shipped — models operating real software on their own, not just chatting about it. Per OpenAI, that capability keeps climbing each release.
What to do with this
- List ten recurring deliverables. Proposals, estimates, month-end summaries, intake write-ups. Mark which are well-specified — known inputs, a fixed format, a clear definition of done. Those are automation candidates. The rest need specification before they need AI.
- Watch work benchmarks, not coding leaderboards. Unless you sell software, a GDPval-style score tells you more about your next automation than any coding chart. Vendor pages will keep leading with these numbers; hold them to the deliverables you actually produce.
- Default to the base tier. Start at $5 and $30, and escalate to Pro only when the standard model demonstrably fails on a task that matters. Nobody should pay six times the rate by default.
- Treat day-one numbers as claims. Every figure above is OpenAI's own, unverified on launch day. Run the model on three of your real deliverables before believing any of it.
If you want a second opinion on which of your deliverables are ready to hand to a model — and which need specifying first — that's a 30-minute conversation.