OpenAI shipped three new voice models today: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper — described by the company as "new realtime voice models in the OpenAI API that can reason, translate, and transcribe speech." All three live in the Realtime API, which makes this a developer release: nothing to download, everything to build on.

The split is the design. GPT-Realtime-2 is the conversational model — OpenAI says "The model can handle harder requests and carry the conversation forward naturally." GPT-Realtime-Translate does live interpretation: "It translates speech from 70+ input languages into 13 output languages while keeping pace with the speaker," per the launch page. GPT-Realtime-Whisper is the transcriber — OpenAI's claim: "It transcribes speech live as the speaker talks."

The pricing is the part worth reading twice. GPT-Realtime-2 is priced in tokens: $32 per million audio input tokens ($0.40 cached) and $64 per million audio output. The other two are priced by the minute — $0.017/min for transcription, $0.034/min for translation. Converted into units a business actually plans in: live transcription runs about $1 per hour of audio, and live translation about $2.

The meter is the story

In most 2–50-person businesses, the highest-value information in the company moves through phone calls — and evaporates on hangup. The price you quoted verbally. The commitment to have a crew out Thursday. The complaint that never made it into the CRM. At $1 per hour, capturing all of that stops being a project budget and becomes a rounding error: a service business that spends 20 hours a week on the phone can transcribe every minute of it for roughly $20 a week.

The transcript itself isn't the win — it's the raw material. Text is searchable, and text is cheap to process: a second, inexpensive model can pull out every quote given, every commitment made, every complaint raised, and file them where someone will act on them. Voice was the last major channel in a small business that didn't produce usable data. If OpenAI's per-minute prices hold, that stopped being true this week.

Translation is the same math pointed at revenue instead of records. OpenAI's claim — 70+ input languages into 13 outputs "while keeping pace with the speaker" — is a market-reach statement in disguise. If your team passes on jobs because a call would have to happen in a language nobody on shift speaks, $2 per hour is now the listed price of not passing. That's a growth number, not an IT number.

The usual day-one caveats apply. Every claim above is OpenAI's own, published on the launch page, and unverified by anyone else as of this writing. "Keeping pace with the speaker" needs testing against your accents, your jargon, and your bad speakerphone audio before it touches a real workflow. And deciding which of the three models fits your operation — if any — is a matching-the-tool-to-the-task question, the same Solution Levels discipline we apply to every engagement. (OpenAI is pushing voice on the consumer side too — that's a separate story, covered in conversation, not turn-taking.)

What to do with this

  • Put a number on your phone hours. Estimate weekly call hours across the team and multiply by $1. Then ask what a searchable record of every call would settle — the next disputed "you said Thursday" gets resolved by search, not memory.
  • Treat language coverage as market sizing. List the calls, customers, or neighborhoods you currently lose to a language barrier. At $2 per hour of live translation, that list is a priced growth experiment — run it through the three-question filter like any other AI idea.
  • Start with transcription, not conversation. The talking model is the hardest of the three to deploy well; the transcriber is the boring one that pays. A transcribe-and-extract pipeline on one call type — inbound quotes, say — is the shape of a first win you can ship in 30 days.

If you want help deciding where voice fits your operation — or whether it does yet — that's a 30-minute conversation.

Source: Advancing voice intelligence with new models in the API — OpenAI, May 7, 2026