Field notes on making AI work
Practical guidance on finding the right use cases, driving adoption, and shipping AI that actually helps — no hype.
GPT-Live: conversation, not turn-taking
OpenAI's new voice models listen while they speak. That resets what customers expect from every automated conversation with your business.
Grok 4.5: the coding agent grows into knowledge work
Cursor and SpaceXAI's new model was trained past software into data, finance, and legal work. The skill that transfers is the one businesses should watch: finishing long tasks with tools.
Claude Sonnet 5: serious AI capability just got cheaper
Anthropic's new mid-tier model runs close to its flagship at a fraction of the price. For small businesses, the economics of automation just shifted again.
Most “AI ideas” aren’t AI use cases
The fastest way to waste a quarter is to build the impressive demo instead of the useful one. Here’s how we separate the two.
Claude Fable 5: the frontier now has a governance story
Anthropic shipped its most capable model, and export controls followed within three days. That timeline is information, not noise.
The AI operators already in your company
Adoption doesn’t start with a rollout plan. It starts with the handful of people already quietly using AI to do their jobs better.
Claude Opus 4.8: agents that finish — and check their work
Anthropic's new model finishes every case on its agent benchmark and is 4x less likely to let its own flaws slide. That changes delegation.
A 30-day path to your first AI win
You don’t need a transformation program to get started. You need one useful result, shipped, that earns the right to the next one.
GPT-Realtime-2: voice AI, billed by the minute
OpenAI's new realtime models put live transcription near $1 an hour and translation near $2. That changes what's worth recording.
GPT-5.5: the benchmark that measures your job
OpenAI says GPT-5.5 scores 84.9% on GDPval — a test built from real work in 44 occupations. Watch that number, not the coding charts.
Claude Opus 4.7: how to read a point release
Opus 4.7 looks like a minor update. It isn't: four tasks no predecessor could solve, and Replit reports same quality at lower cost.
Composer 2: the price floor for agentic coding just fell
Cursor's new model prices frontier coding at $0.50/$2.50 per million tokens — a tenth of flagship rates. The build-vs-buy math just moved.
GPT-5.4: the model that operates your software
OpenAI says GPT-5.4 finishes 75% of real computer tasks on its own, up from 47%. AI that drives your existing tools just got plausible.
Claude Sonnet 4.6: the default tier just caught up
Anthropic says its default model now matches its flagship on office work — same $3/$15 price. The wait for 'cheaper' keeps shrinking.
Composer 1.5: fast and good enough beats slow and perfect
Cursor's new coding model thinks hard only when the problem deserves it. The same effort-matching should govern which models you pay for.
Claude Opus 4.6 can hold your whole business in its head
Anthropic's new flagship reads a million tokens at once and triaged a day's work queue on its own. That's a dispatcher, not a chatbot.
Ready to put this into practice?
Let's talk about where AI fits in your business.