Big AI initiatives stall because they try to boil the ocean. The teams that succeed do the opposite: they ship one useful thing quickly, then build on the trust it earns.

This is the playbook we run for that first month. To make it concrete, we’ll follow an example team through it: a twelve-person specialty contractor whose estimates take three days to go out, mostly because turning site-visit notes and photos into a written quote is nobody’s favorite job. The details are illustrative — swap in your own slow, repeated, language-heavy task and the shape holds.

Before day one

Three things need to exist, or the month stalls in week two:

  • A decision-maker who will spend one hour a week. Not to manage the work — to remove blockers and say "yes" fast when the pilot needs something.
  • One operator. The person who already experiments with AI on their own. If you don’t know who that is, here’s how to find them.
  • Access to where the work actually lives. If the notes are in a clipboard app and the quotes are in a spreadsheet, the pilot needs both. Half of first-month friction is permissions, not models.

Week 1 — Find the target

Interview the people doing the work. Look for a repeated, language-heavy task with an obvious cost when it’s slow or skipped. Pick one. Write down what "better" would mean in plain numbers.

At the example contractor, the interviews surface three candidates: writing estimates, chasing overdue invoices, and answering "where’s my crew?" calls. Estimates win — they happen ten times a week, every day of delay costs jobs to faster competitors, and the raw material (notes, photos, past quotes) already exists. The plain-numbers target gets written on an actual whiteboard: "Quote turnaround: 3 days → same day, without losing accuracy."

If several ideas seem plausible and you can’t choose, that’s a sign the list needs the three-question filter before the calendar starts.

Week 2 — Build the smallest version

Solve the narrow problem, not the general one. Use existing tools where they’re good enough; build only what you have to. The bar is "measurably better than today," not "finished forever."

For the contractor, the smallest version is deliberately unglamorous: a drafting assistant that takes the estimator’s site notes and photos and produces a first-draft quote in the company’s format, using the last two years of quotes as reference. It doesn’t price anything on its own — pricing judgment stays with the estimator. It doesn’t integrate with accounting yet. It writes the draft; a human finishes it. That restraint is the point: every feature you skip in week two is a failure mode you don’t have to debug in week three.

Week 3 — Put it in front of a real user

Not a demo — a real person doing real work. Watch where it helps and where it gets in the way. Fix the friction. This is where most of the value is found.

Week three at the contractor is humbling, which means it’s working. The estimator ignores the tool for two days (nobody budgeted time to try it — the decision-maker’s one weekly hour fixes that). The drafts mangle a material name that appears in the notes as shorthand. Photos help less than expected; the past-quote library helps far more than expected. By Friday the estimator has stopped saying "the AI" and started saying "my draft" — the single most reliable adoption signal we know.

Two rules keep this week honest. First, fix friction the user actually hits, not friction you imagined. Second, resist every "while we’re at it" — new scope goes on a list for later, not into the pilot.

Week 4 — Measure and decide

Compare against the number you wrote down in week one. If it moved, you have a win worth expanding and a team that now believes. If it didn’t, you’ve learned something cheaply and can redirect.

In our example, the whiteboard number moves: quote turnaround goes from three days to same-day on most estimates — a draft is ready before the truck is back at the shop, and the estimator’s hour of writing becomes fifteen minutes of editing. Just as important is what didn’t change: accuracy stayed put, because the human judgment never left the loop. That’s the story the decision-maker tells in the next all-hands, and it lands better than any vendor deck because the person who proved it sits three desks away.

Measure the boring way: the same number, before and after, no cherry-picking. One honest metric beats five flattering ones — the flattering ones get audited eventually, usually in the budget meeting where you ask for the second project.

What goes wrong (and how to avoid it)

  • The scope quietly generalizes. "Draft the quote" becomes "handle all customer communication." Kill this in week two; the general version is a different, much harder project.
  • The task has no owner. If nobody does the task daily, nobody will notice whether the pilot helps. Pick a task with a name attached.
  • The demo-to-daily gap. Working once in a meeting is not working every morning at 7am. Week three exists precisely to cross that gap — don’t skip it to hit a demo date.
  • Vanity metrics. "The team loves it" is not a number. If week one’s target wasn’t written down, week four becomes a vibes discussion.

The point of the month

The point of the first 30 days isn’t to finish. It’s to prove the approach works in your business — with your data, your people, and your constraints — and to make the next step obvious. The second project starts with something the first one didn’t have: a team that has seen this work.

This playbook is what our Point Solutions engagement runs end to end — scoped in week one, live by week four. If you have a candidate task in mind, tell us about it; the first conversation is 30 minutes, and you’ll leave with an honest read on whether it can carry the month.