Skip to content

What Good Looks Like, and the Three Questions to Ask

Before you start

Prerequisite: Lessons 01-02, the harness-vs-model distinction and the five warning signs. This lesson assumes you can already spot a weak project; it adds the positive checklist for a strong one. After this lesson, you can: recognize the staged-autonomy rollout pattern, calculate why a lower automation rate can be the safer business choice, and carry three questions into any AI funding conversation.

Funding red flags is only half the job. This lesson is the positive checklist: the shape of an AI agent project actually worth backing, and the three questions worth carrying into every AI conversation you have after this one.

Autonomy is earned in stages

The proven rollout pattern has three steps, not one leap. Stage one is human-review mode: every AI response gets checked by a person before it goes out, and the team tracks the override rate. Stage two is low-risk autonomy: the system handles only the reversible, low-stakes band on its own. Stage three is expanded scope: the team widens what's automated only as its evaluation results earn it.

The pitch "full autonomy in weeks" is itself a red flag, not a selling point.

The most mature voice in the room says start smaller

Anthropic's own public guidance on agentic systems: "find the simplest solution possible... this might mean not building agentic systems at all." A vendor confident enough to tell you that is more trustworthy than one promising magic by Friday.

Autonomy earned in stages, not granted on day one

Staff for judgment, not just models

The picture most people carry in their heads is a room full of AI and machine learning engineers. That picture is incomplete. Those engineers are necessary (they build the engine) but a working AI project is mostly other people: evaluation and QA staff who define and guard what "right" means technically, domain experts who decide what "right" means in the business, operations staff who keep the system alive at three in the morning, and governance and compliance staff who own the liability question from the previous lesson.

IBM's Cost of a Data Breach research found that 63% of breached organizations had no AI governance policy at all. That's not a statistic to file away. It's close to the job description of whoever in the room ends up owning AI.

Buy the model, own the harness decisions

The model and the platform are increasingly things you buy. The harness decisions specific to your risk and your data (evaluation criteria, escalation rules, who owns liability) are not outsourceable. That ownership sits with whoever holds the AI governance role, whether the title is Chief AI Officer or just "the person who ended up owning this."

Measure quality, not motion

Some metrics measure how much a system moved. Others measure whether the customer actually walked away helped. The second set is what's worth reporting to a board; the first is vanity dressed as progress.

Motion (vanity)Value (what counts)
Percent automatedResolution quality
Deflection rateRepeat-contact rate
Tokens usedCost per successful task
"Agents deployed"CSAT on the automated slice

The arithmetic makes the tradeoff concrete. Take 100 customers. A team chasing automation rate runs 75% automated at 90% accuracy: 67 customers correctly resolved, and 8 wrong answers reaching real customers. A more disciplined team runs 50% automated at 98% accuracy: 49 customers correctly resolved, and the other 50 cleanly escalated to a human, leaving only 1 wrong answer in front of a customer. The "lower" automation rate produces roughly a quarter of the customer-facing errors. That's the number that compounds, because every wrong answer that reaches a customer is a support ticket, a lost account, or, per the previous lesson, a liability event.

Push quality up first

Get the automated slice accurate before you touch how much is automated.

Widen scope only as evals earn it

Expand what the system handles automatically only after the evaluation numbers support it, not on a calendar deadline.

Report cost per successful task

Not tokens, not deflection rate. The unit that tells you whether you're funding a business or a demo.

Klarna is the company that lived this tradeoff in public and corrected it. Its CEO, Sebastian Siemiatkowski, said in May 2025: "cost was too predominant a factor... what you end up with is lower quality." The correction that followed is the part worth copying: Klarna didn't rip the AI out. It kept the system carrying the bulk of the volume and added humans back specifically for disputes, fraud, and hardship cases, with a guaranteed commitment that a customer can always reach one.

The specific headcount and dollar figures Klarna has cited publicly for this reversal are the company's own projections and press statements, not independently audited numbers, worth treating as directional rather than as fact. What's solid is the pattern: the public admission of a quality problem, and the corrective action that followed it.

The headlines are all the same shape

Air Canada, February 2024: a chatbot's invented refund policy ruled binding. McDonald's and IBM, June 2024: a drive-through AI pulled after three years, following widely shared incidents involving wildly wrong orders. Klarna, May 2025: the public reversal above. Taco Bell, August 2025: 18,000 cups of water crashing the order system, covered in the first lesson of this course. Munich v. Google, May 2026: a company held directly liable for its AI's own false answers.

None of these were unserious teams. These are some of the best-resourced companies operating today. In each of these cases where the underlying facts are fully settled (Air Canada, McDonald's/IBM, Klarna, Taco Bell), the failure traces to the harness, not the model: a missing guardrail, a broken handoff, no liability plan, no cost cap. The Munich ruling is the newest of the five and still developing as this course is written, so it's included as a same-shaped warning sign, not as a fully closed case. The lesson isn't "don't ship AI agents." It's "fund the harness, not the demo."

Three questions for your next AI conversation

Judge the harness, not the demo. A demo proves a handful of steps on a good day. Ask what's built around the model: evaluation, escalation, guardrails, grounding, monitoring.

Start where being wrong is cheap. Fund low-risk, reversible, well-measured pilots in human-review mode first. Treat "full autonomy in weeks" as a tell, not a promise.

Measure quality, not motion. Resolution quality, repeat-contact rate, cost per successful task, and who's liable when it's wrong. Never tokens, never deflection rate alone.

Quick check — Across 100 customers, which setup leaves fewer wrong answers in front of real customers: 75% automated at 90% accuracy, or 50% automated at 98% accuracy?

The next lesson moves from the executive framing in these three lessons into the discipline itself: why harness engineering exists as a named practice, and what it actually involves to build one.

Continue to Lesson 04

Why harness engineering exists as its own discipline, and what separates it from simply "using a better model."

Have a question about this lesson?

Reply here and it goes straight to Rod. Same as replying to one of his emails.