OpenAI released GPT-5.4 on March 5 in three variants: GPT-5.4, GPT-5.4 Thinking, and GPT-5.4 Pro. The company describes it as "OpenAI's most capable and efficient frontier model for professional work, with state-of-the-art coding, computer use, tool search, and 1M-token context." API access is immediate — gpt-5.4 at $2.50 per million tokens, gpt-5.4-pro at $30 — while ChatGPT Plus, Team, and Pro accounts get it in a gradual rollout.

The launch numbers are OpenAI's own. On professional work, the company says GPT-5.4 "achieves a new state of the art, matching or exceeding industry professionals in 83.0% of comparisons." On web research, OpenAI says it "achieves a leading 89.3% accuracy on BrowseComp." And the number this note is about: per OpenAI, the model "achieves a state-of-the-art 75.0% success rate on OSWorld-Verified, far exceeding GPT-5.2's 47.3%." OSWorld measures whether a model can operate real computer environments — actual apps, actual windows — on its own.

The rest of the sheet: the release folds in GPT-5.3-Codex's coding capabilities, OpenAI calls it "our most token efficient reasoning model yet," and the Thinking variant "can provide an upfront plan of its thinking, allowing adjustments mid-response." All of it is vendor-reported until it survives contact with your own work.

Operating a computer, not just chatting

For three years, the practical objection to AI automation in a small company has been the same: the work lives in a browser tab, a spreadsheet, and a CRM, and the model can't touch any of them. A chat answer still needed a person to carry it into the real system. That's why so many AI projects shipped as a text box beside the actual work instead of inside it.

OSWorld tests exactly that gap: hand a model a real desktop and a task, then see whether it finishes without help. At 47.3% — GPT-5.2's score, per OpenAI — it fails too often to trust with anything a customer sees. At a claimed 75.0%, it looks more like a new hire in week two: right three times out of four, usable with someone checking the output. That's a 28-point jump in one generation, and it moves "AI that uses your existing tools" from demo to plausible.

The business translation matters more than the benchmark. Automation no longer has to start with replacing your software. The model is learning to work your tools where they are — the same browser, the same spreadsheet, the same CRM your team logs into today. That inverts the usual cost math: in most small-company automation projects, the expensive part was never the AI. It was the migration.

The second thing worth noticing is what comes out the other end. OpenAI says GPT-5.4 handles "long-horizon deliverables such as slide decks and financial models" — finished artifacts, not fragments a person still has to assemble. A reviewable draft deck is a different product from twelve bullet points about one. Where OpenAI took this next is its own story — see GPT-5.5: the benchmark that measures your job.

What to do with this

  • Inventory the work your team does inside software, not about it. Copying fields between two systems, pulling numbers from a portal into a spreadsheet, updating records after every call. A year ago those were dead ends because no model could drive the screen. They're candidates now.
  • Pilot with a reviewer, not autonomy. 75% per OpenAI still means one failure in four. Start where a human already checks the output before it goes anywhere, and hold the model to the same standard as a new employee: supervised until proven.
  • Don't budget for new software. The point of this release is the opposite — your existing stack is the surface the model works on. If a vendor's pitch starts with a migration, ask why. Matching the tool to the task is how we scope every engagement.
  • Ship one boring task in 30 days. Pick the dullest candidate on your inventory and run it through the 30-day playbook. Small, checked, measured.

If you want help picking the first computer task worth handing to a model, that's a 30-minute conversation.

Source: Introducing GPT-5.4 — OpenAI, March 5, 2026