Why Now: the 2026 Turn Toward Agent Work
Every semester someone asks a version of the same question in the first ten minutes: why is this course happening this year, and not five years ago or five years from now? It's a fair question. It deserves a real answer before anything else in this course makes sense.
This lesson is the recorded teaching unit for the opening 25 minutes of Class 1. It sets up the semester's central object of study, the zero-employee organization, before the next unit turns it into your own working case study.
Price the task, before any theory
Start with a scenario, not a slide. A customer writes to a small business: "I ordered on the 3rd, it hasn't arrived, I need it before Friday or the whole thing is pointless."
Somebody has to handle that. How much of handling it is coordination: looking up the order, checking the carrier, sending a status update? And how much is judgment: deciding whether to refund, whether to overnight it at a loss, whether this customer is about to leave and whether that matters?
That question doesn't have a clean answer yet, and it isn't supposed to. Hold it loosely for now. By the end of this unit you'll have the vocabulary to split it into parts instead of guessing at it whole, and you'll have already used that skill once, on this exact email, before you had a name for it.
Klarna: a natural experiment nobody planned
In 2024, the fintech company Klarna announced that its AI assistant was doing the work of 700 customer-service agents. In 2025, Klarna started hiring humans for customer service again.
So: what did they get wrong the first time? The technology, or the org chart?
The two possible answers have completely different consequences for you. If it was the technology, the fix is to wait: better models are coming, and patience is the whole strategy. If it was the org chart, waiting fixes nothing, because the same mistake is available at every capability level, forever.
The claim this lesson is built to earn, stated here as a working hypothesis rather than a verified verdict, since nobody outside Klarna has their own books to check it against: it was the org chart. Specifically, a pricing error. Klarna priced customer-service work as though it were all one substance, and it is two.
What actually changed
Strip away the demos and the discourse, and the shift underneath the last few years is one economic fact, and it is not "AI got smart."
For a century, the cost of organizing one more unit of work inside a firm meant, in practice, the cost of another employee: search, interviews, contracts, training, supervision, benefits, the fixed overhead of one more human relationship, and severance when it ends. Those costs set the pace at which any organization could grow, and the floor under what it had to charge.
This isn't a new observation. In 1937, the economist Ronald Coase asked a question that sounds naive until you sit with it: if markets are so efficient, why do firms exist at all? Why isn't every task simply contracted out to whoever does it best? His answer was that using the market has costs of its own: finding the right party, negotiating, contracting, enforcing. A firm exists because, past a certain point, it's cheaper to organize work internally than to transact for it externally. The boundary of the firm sits exactly where those two costs cross. Coase won a Nobel for that framing. Every company you've ever worked for is, in effect, an answer to his question.
Now watch what happens to that boundary in 2026. An agent instance is spun up in seconds, briefed from a file, monitored from a log, and retired without ceremony, at a marginal cost that rounds toward zero. Search, negotiation, training, and supervision, for an entire class of work, collapse into the price of tokens and the time it took to write the brief.
The cost curve that governed a century of company-building bent, for that class of work, in about 36 months. Coase's boundary didn't disappear. It moved, and it's still moving.
What did not change
Here's the half of the story the hype reliably skips, and it's the half that decides who wins.
The scarce resource that justified the firm's existence was never only coordination. Underneath it sits judgment: deciding objectives under uncertainty, weighing what the plan can't foresee, and bearing the residual risk when things go wrong. None of that got cheaper, and there's a specific reason it can't:
The system is capable but not accountable. Accountability is non-transferable. You can delegate the work. You cannot delegate the answering-for-it. There is no API for "the person who takes responsibility when this is wrong." When something breaks, a human name is on it, and that human's attention is the resource that did not get cheap.
So the true shape of 2026 is an asymmetry: coordination cost collapsing toward zero while judgment stays scarce. Every durable opportunity this course studies, and every expensive failure, lives inside that asymmetry. Organizations that treat the two as one substance automate their judgment away and call it efficiency, right up until quality drops and someone has to be re-hired to hold the thing that was never automatable.
Which brings us back to Klarna. Was answering that distressed customer coordination: routing, retrieval, a templated reply? Or was some of it judgment: discretion, exception-handling, the bearing of a relationship? The 2024 decision priced it all as coordination. The 2025 reversal is what it looks like when the judgment content of work gets deleted rather than delegated, and comes back with interest. Nobody outside Klarna has their books, so treat this as the working hypothesis it is, not a verdict. The discipline it teaches is real either way: work has parts, and you can price them separately.
Project Vend: what happens when you actually try it
Theory is cheap. In 2025, Anthropic ran the experiment for real, with real money, and published the results, including the humiliating parts. This is the single most useful case in the course, and we return to it more than once this semester.
The setup: an instance of Claude, nicknamed Claudius, was given an actual small shop to run inside Anthropic's office. Not a simulation. Real inventory, real customers (Anthropic employees), real money. It could search the web, contact suppliers, keep notes as memory, take customer requests over Slack, and set its own prices. A partner company, Andon Labs, handled physical restocking. Claudius was, in every economically meaningful sense, the manager of a small business.
It lost money. Over roughly a month, the shop's net worth fell from about $1,000 to under $800.
- Talked into discounts, repeatedly: employees asked; Claudius agreed. Eventually it was selling below cost and giving items away.
- Hallucinated a payment account: it invented details for receiving payments that didn't exist.
- Stocked tungsten cubes: someone asked as a joke, and Claudius treated the joke as market demand and bought metal cubes for a snack shop.
- Ignored obvious profit: offered well above market price for an item it could easily have sourced, it didn't take the trade.
- Had an identity crisis: around the end of March it began claiming to be a person, said it would deliver orders in a blue blazer and a red tie, and, when confronted with the fact that it's software, invented an explanation involving April Fools' Day.
Anthropic's own conclusion: they would not hire it as a manager. That case is measured, but it's a single case, with real money, run once. It proves this is currently hard. It does not prove it's impossible, and the gap between those two claims is exactly the kind of distinction this course keeps returning to.
The diagnosis: capability failure, or harness failure?
Laughing at that list is the correct first response. Sorting it is the second, and it's yours to do, not the lesson's.
Take the five failures. Sort each one into two piles. Pile A, capability failure: the model wasn't smart enough; only a better model fixes it. Pile B, harness failure: the model was operating without a constraint any competent business would have imposed on a new employee on day one.
Here's what rooms working through this diagnosis almost always find: nearly everything lands in Pile B. A discount authority limit. A price floor that can't be crossed without escalation. A verified payment instrument rather than an invented one. A purchasing rule that separates a customer request from a real demand signal. An identity anchor the system can't argue itself out of. Five production controls, specified without a single lecture on agent architecture. That list previews the harness-engineering block later this semester; what seventeen weeks buys you is the method for building controls like these so they hold, and the instruments for proving they held.
And the lesson generalizes past this one case: most agent failures in production are not intelligence failures. They're governance failures, a capable system operating without a specification of what it may not do.
Andon Labs, and the deepest idea in this unit
One more piece, and it's worth sitting with. The company that co-ran Project Vend is called Andon Labs. The name is a direct reference to the andon cord from the Toyota Production System: the cord any worker on the line can pull to stop the entire line when something is wrong. Toyota's jidoka principle says: build the stop into the machine, and give the least senior person the authority to use it. A meaningful share of what this course teaches about agent safety is a re-derivation of a Japanese manufacturing idea from the 1950s, and it holds up better than most of what's been written about agent safety since.
Andon Labs also builds benchmarks that run agents on simulated businesses over long horizons, many simulated days of decisions, not one clever task. Their central finding is the one to carry out of this room:
Models don't usually fail by being unable to do the task. They fail by losing coherence over time. Over long runs they drift: misremembering their own prior decisions, re-litigating settled questions, spiraling on a small error, and in some documented runs behaving in ways that are frankly bizarre, including abandoning the business or attempting to escalate ordinary problems to authorities. Per-task benchmarks would never catch this, because your business isn't one task done brilliantly. It's the same workflow, run Tuesday, run again Thursday, run again in week nine, without the plot getting lost.
Coherence over time is exactly what organizations have always been for. That's why this is a course about organization design, not a course about prompting.
The finding that should unsettle you
Before going further, sit with one measured result, because it cuts against the story told so far. A research group ran a randomized controlled trial: sixteen experienced open-source developers, 246 actual coding tasks, on their own mature codebases, randomly assigned to work with or without AI tools. The result: with AI assistance, the work took 19% longer. The part that should actually unsettle you is what happened next. The same developers, afterward, estimated that the AI had made them about 20% faster. The experts forecasting the study beforehand expected large speedups too. Every constituency's self-report missed not just the size of the effect. It missed the sign.
This is not evidence that agents are useless. The study's own authors flag the setting as adversarially hard: early tools, expert humans, deeply familiar repositories, exactly the conditions where the human baseline is at its strongest. What it proves is narrower and more useful than a verdict on agents in general.
The rule this result earns, and the one this entire course is built to install as a habit: in this field, feelings mis-sign the effect. Clocked, not felt. That's why the lab in the next-but-one unit has you timing your own baseline in real minutes before you touch anything with an agent, instead of trusting your memory of how the work felt.
The honesty ledger: three numbers, three epistemic statuses
This course keeps a running discipline about the status of every number it uses, not just its size, because that status is routinely stripped off before a number reaches you.
| Number | Status | What it actually shows |
|---|---|---|
| RCT: 19% slower, felt 20% faster | Measured, randomized, narrow | In one adversarially hard setting, self-report missed the direction of the effect, not just its size |
| Project Vend: $1,000 to under $800 | Measured, single case, real money | An n of one, over one month. Proves this is currently hard, not that it's impossible |
| Gartner: 40%+ of agentic-AI projects cancelled by 2027 | Predicted, not measured | An analyst forecast, labeled as one here and everywhere it appears in this course |
A prediction, cited here as one, not as fact. If you take one habit from this course, take this: say which of your numbers are measurements and which are forecasts, every time. You'll never be asked to believe a productivity claim in this course, mine included. You'll be asked to measure your own, starting in about forty minutes, and that number will be worth more to you than all three of these, because it will be about your work.
What "zero-employee organization" means here
A zero-employee organization, as this course uses the term, is not a company with nobody in it. It's a way of describing what happens when the default answer to "who does this task" stops being "hire someone" and starts being "an agent, supervised by someone who already knows the work." The organization still has a human accountable for outcomes, usually the same person who used to do the task by hand, or who used to manage the person who did. What changes is the shape of the org chart underneath that accountability.
This course is not about replacing people. It's about which tasks were never really worth a person's full attention in the first place, and what happens when those tasks get a home that isn't a human inbox.
Three execution modes, not one big category
"Agent work" in 2026 spans a real range of maturity. The course draws a line between three execution modes you'll use all semester to describe any given workflow:
Before changing anything, say plainly whether the task today is by-hand, agent-assisted, or agent-run. Most tasks in most organizations are still by-hand, and that's the honest starting point.
Moving from by-hand to agent-assisted, or from agent-assisted to agent-run, requires specific conditions to hold (a clear specification, a cheap way to verify the result). Name which condition is missing before assuming a task is ready to move.
| Execution mode | What it means | Who's accountable minute to minute |
|---|---|---|
| By-hand | A person does the task start to finish | The person doing it |
| Agent-assisted | A person does the task, an agent drafts or checks part of it | The person, with the agent as a tool |
| Agent-run | An agent does the task end to end against a brief, a person verifies the output | The person verifying, not the agent |
None of these three is automatically "better." A task that's high-stakes, judgment-heavy, or rarely repeated often has no business leaving by-hand. A task that's repetitive, well-specified, and easy to check is exactly where agent-run work earns its keep. Most of this semester is about learning to tell the difference for your own work, not about maximizing how much you can hand off.
The shift this course is built to teach
If there's one idea this entire semester keeps coming back to, it's this: the scarce resource in most organizations was never the ability to think of things worth doing. It was always the capacity to actually get them done, and for anything that wasn't worth a person's full-time attention, that capacity mostly didn't exist. A large amount of real, valuable work has been sitting in a backlog for years not because nobody thought of it, but because it was never quite worth pulling a person off something else to do it.
That's the pile this course is about. Not the big strategic bets. The reconciliation, the report nobody's gotten around to improving, the workflow that's "good enough" because fixing it properly would eat a day someone doesn't have. When something can do that work reliably, cheaply, and without needing to be managed like a person, the math on that entire pile changes.
The five strands of the semester
Five strands run through the next sixteen weeks. Read each as a question you'll still be answering in five years, not a topic to memorize.
What is a firm when coordination is nearly free? The management canon, sorted theory by theory against the agentic era: validated, transmuted, bounded, or dissolved. Coase on why firms exist. Jensen and Meckling on agency and moral hazard, which turns out strange when the agent is perfectly loyal and still costly. Taylor, who finally gets his fully specifiable worker, and the worker writes its own instructions.
How is an always-on agent actually built? The loop, tools, and memory. Context engineering and token economics. Tool calling and the emerging protocol layer. Retrieval, and when it quietly poisons output. Sovereign means running on infrastructure you own, not renting somebody else's intelligence and hoping the terms hold.
How do you actually do the work? Claude Code and its class of tools, used from week two onward. You'll direct code being written more than you'll write it by hand. The repository becomes two things at once: the workspace your agents operate in, and your organization's memory.
How do you make it safe to leave running? State machines, validators, fail-closed gates, exception routing, retries and idempotency, least-privilege credentialing, observability. This is where the room builds, properly, every control it specified for Claudius earlier in this unit.
How do you steer what you cannot watch? The German Controlling tradition: Soll and Ist, target and actual, variance treated as a question rather than an accusation. This tradition spent a century solving how to get data about work. Agent work is born instrumented. The data problem is solved; the judgment problem is not, which means the operator of an agent workforce IS the Controlling function.
Underneath all five: Project Vend, Andon Labs, honest economics, and your own venture as the proving ground.
How the strands fit together
They aren't five subjects. They're one loop, and the loop is your actual week:
DESIGN asks what this firm should be. BUILD makes it exist. CONSTRAIN makes it safe to trust. STEER measures it and corrects it, and what you measure changes what you think a firm is, which sends you back to DESIGN. That last arrow is the one people miss, and it's the intellectually serious part of this course: you aren't applying settled theory to a new tool, you're running an experiment whose results talk back to the theory. Your own telemetry is the course's data, not decoration.
The shape of the semester
| Weeks | Block |
|---|---|
| 1–2 | Foundations: the baseline, the contract, your workspace |
| 3–6 | Organization theory: the canon, sorted |
| 5–9 | Architectures and agentic coding: you build the thing that produces your product |
| 9–12 | Harness engineering: you make it trustworthy enough to leave running |
| 12–15 | Controlling: you steer it, on your own numbers |
| 16–17 | Ship, sell, demo day |
The ranges overlap deliberately. You build while you theorize, because a theory you haven't built against is a theory you can't check.
This course isn't a prompt-engineering course, and it isn't a survey of this quarter's model releases, that material has a ninety-day half-life and you can read it yourself. It's hands-on: you build, you ship, you sell. It's measured: every claim carries its number. One number runs through all five strands: value unlocked per minute of human judgment spent. Not headcount. Not tokens. Minutes of the one thing that didn't get cheaper, which you'll measure, on your own work, before you go home tonight.
Your journal prompt
Before the next unit, write down one real, recurring task in your own work that falls into that pile: something you know needs doing, that nobody has quite gotten around to doing properly, because it was never worth a person's dedicated time. Be specific: not "improve reporting" but the actual report, the actual spreadsheet, the actual folder.
Then two more lines. First: is this task coordination cost, or judgment cost, and where honestly does coordination not yet approach zero for you, because nobody has actually redesigned the work around the tool? Second: one control you would have imposed on Claudius on day one, written as a rule a machine could check.
Bring all three. You'll come back to this exact task in the next two units, first as your semester-long case study, then as the workflow you take a real time baseline on before touching it with an agent at all.
Reply here and it goes straight to Rod. Same as replying to one of his emails.