About this episode
Your team lead runs an AI pilot, the agents spend millions of tokens in a week, and nobody can point to one thing the team used. The agents started from the files alone, and the people who joined the team this year still guess at how decisions get made. Your agents and your people are missing the same clarity about what matters, what is locked, where to stop and who decides. That clarity is the bottleneck on your people and on your machines.
The ninth essay in the Where Signal Gets Lost arc, as an operating brief. Context Equalization, the intervention in the Cognitive Translation Protocol where a team makes its context clear for its people, turns out to be what the AI agents your organization runs need too. The same ask goes to Claude Fable 5.1 twice on one night, once bare and once with the team's context. With the context, the model's own estimate of the work falls from roughly 12 to 17 hours to roughly three to four, and it is just as careful in both runs. The research behind rising AI token costs, what the makers of these models tell their own developers, and the design change your organization can make to give your people and your agents the same clarity about how the work gets decided. Episode 9 of the arc.
Full transcript The episode as text · lightly cleaned for reading
Somewhere in your organization, a team lead is running an AI pilot this week. On Friday, the bill arrives, and it is millions of tokens.
Here is what those tokens bought. Over one week, the agents read three years of that team's files. They built work breakdown structures for four separate workstreams. They ran a dozen more agents to verify the work and to argue against it. And the agent running all of it finished the week sure of what it had produced.
Then you ask your team lead the plain question. What are we actually using out of all that? And your lead cannot point to one thing.
I want to stay on that Friday, because what your lead does next is what this episode is about.
Your lead sits down and writes out everything the next run will need to know. What this piece of work is actually for. Which numbers are settled and which ones are guesses. What nobody is allowed to touch. Where to stop and come ask. What to drop if the deadline moves up. And who makes the final call.
That takes your lead about a week. At the end of it there is one page. Six lines on it, in plain words, about one piece of work. And that page turns out to be the first written description anyone on that team has ever made of how the team decides anything.
Then your lead reads it back, and sees something nobody was looking for. Two people joined that team in April. They have been guessing at every one of those six lines for five months. Everyone who had been there longer knew all of it, and never said it out loud, because until the agents showed up, nobody ever had to.
So here is the thing I want you to carry out of this episode. Every thinker in your organization, human or machine, runs on what somebody wrote down.
That is why I say communication in your organization is now about clarity, and specifically the written kind. And that written clarity is your bottleneck. It is the bottleneck on your people. And it is now the bottleneck on your machines too.
You are not the only leader hitting it. In McKinsey's State of AI survey this year, about one in five of the people they surveyed said AI operating costs were holding back how much AI their organization used. Token costs are part of that. And most of those same people still planned to spend more. Now hold that next to the return. The share of them who said AI had moved their organization's earnings at all was just under four in ten. That is about where it was the year before. All of it is people reporting on their own organizations, not audited numbers.
So the spending climbs and the reported return stays where it was. I read that gap as the same thing your team lead hit on that Friday. Nobody had written down what that work was for. So the agents filled every gap in their instructions with work, and you paid for all of it.
I want to be straight about that link, because it is mine. McKinsey counted the spend and counted the return. Putting the two together and calling missing written context the cause is my reading, not theirs.
And before this starts to sound like a complaint about your team, it is not one. Your organization ran on unwritten context for a good reason. Anyone who had been there a while picked it up by being there. In the meetings. In the hallway. At the next desk. For years that was enough, and nobody had any reason to write it down.
The agents got none of that. They started from the files alone. And the two people who joined in April sat in most of those same meetings, and they still had to guess.
So let me give you the six lines, because they are the practical part of this episode and each one is a sentence you could write today.
The first line says what matters here, and what is not being decided. The second says what is known, and separately, what is only assumed. The third says what is locked, meaning what nobody touches. The fourth says where to stop and check with a person. The fifth says what to drop if time runs short. And the sixth says who decides.
That is the operating page. Six lines, one piece of work, plain words, no template.
Two of those six are the ones people skip. Where to stop and check, and what to drop if time runs short. Those two are the lines that change what an agent costs you.
There is one more step, and it is what makes the page worth the week it takes. Once the six lines are written, read them out with as many different kinds of thinkers as the room holds. Every one of those people catches a gap the others missed. That is judgment, and an agent cannot supply it.
I have written about this page before in this series, under the name Context Equalization. It is one of the interventions in my Cognitive Translation Protocol, which is my framework for where communication breaks down between people who think differently. The stop line and the drop line are what this week adds to it.
Now here is the part that surprised me.
Your people already think in different ways. You have known that for years, and The Neurodiversity Movement is the work that put a name to it. What is new is that your organization now runs frontier AI models beside your people, and those models are not the same as each other either. I read that as a machine version of the same variation you already have in the people on your team.
A frontier model is an enormously capable system. And it starts every single task with nothing except the prompt somebody typed and the files somebody handed it. Around the model there is a harness, which is just the software that gives the model its tools and its files and lets it keep working over time. I use Anthropic's Claude Code as mine. The harness shapes what a model does with a one-line ask.
Here is the part I think leaders have backwards. The most capable systems you can buy are the ones that need clarity and context the most. A model only ever gets what somebody wrote down. And when nobody has written the six lines, the model still has to act.
So it hits a line nobody wrote, something like what is locked, or where to stop, and it fills that gap itself. What it puts there depends on how it was trained.
One kind of training taught a model to answer with confidence anyway. OpenAI's own researchers wrote in September 2025 that the standard way models are trained and scored rewards a confident guess over an admission of uncertainty. Under right-or-wrong scoring, saying I do not know is the worst thing a model can say. So a model that meets a gap learns to fill it. That is an argument they make using their own numbers, not an experiment they ran.
The other kind of training taught a model to keep working on the gap instead. Inside an agent, you see that habit as extra rounds and extra helper agents. Fan and colleagues found in 2025 that reasoning models produced two to four times more tokens on a math question when one premise had been removed from it.
I see this at my own desk. I run this site on Claude Fable 5.1, in Anthropic's Claude Code harness. Give it a one-line ask, and I get back a plan with more workstreams and more subagents than the work actually needs. Anthropic reported in June 2025, from its own data, that an agent uses about four times the tokens of a chat. A system of several agents working together uses about fifteen times.
So put those two kinds of training at opposite ends. At one end, a model trained to guess. It fills the missing line with a confident answer. At the other end, a model trained to reason and act. It fills that same missing line with work. More tokens, more rounds, more helper agents, and a bill you can see. Those two behaviors look nothing alike. Underneath both of them sits one cause, which is the same missing context. And there is one fix at both ends. Write it down, once, for everyone.
And your people have needed those same six lines all along. In an earlier essay I called that cost the translation tax, which is the hours a person spends working out what an ask actually meant. Nobody ever counted those hours. The token bill is that same cost with a number finally attached to it.
How much context each person needs in writing differs, because the people on your team are wired differently. I treat that as variation to design for, not a problem to fix. Here is the premise I start all of my work from: different nervous systems produce different cognitive architectures, shaping thinking, sensing, and processing across mind and body, not broken brains.
I wrote a rule about this into my Cognitive Translation Protocol years before agents existed, and it was about people. It says, "Communication breakdowns happen at the interface between different cognitive systems, not inside either one. The fix is structural, not personal." In plain words, the problem is not inside the person who spoke or the person who listened. It sits in the space between them, and you fix that space by designing it.
With an agent, that space is the only thing there is. The only thing between your team and an agent is what your team wrote down. So the pilot you are running is the first place you can watch that rule work out in the open.
Your agent needs the operating page. So do the two people who joined in April.
So what does one page of six lines actually change? I ran a test to find out. Both replies are printed word for word in the essay, so here I will tell you what happened.
I took one ask and sent it to Claude Fable 5.1 twice, on the same night, in the same harness, with nothing else changed. The ask was six words. Prepare the pricing review for Thursday's leadership meeting. That was the whole of the first prompt.
The second prompt was those same six words with an operating page underneath them. The page said one decision gets made on Thursday, hold list price through the fourth quarter or move it, and nothing else is being decided. It said which numbers were settled. It said, separately, which ones were only assumptions. It said what was locked, including that no customer-facing number leaves the team before the VP has seen it. It said stop and check before two hours. It said drop the competitor teardown if time runs short. And it said the VP of Sales decides, on Thursday.
Here is what came back without the page. Claude Fable 5.1 planned to look around first and ask its questions afterwards. It drafted one batched message to the team lead, with a default answer attached to every question. And it said that if nobody answered by Tuesday end of day, it would go ahead on those defaults. Then it handed the work to three helper agents, kept the write-up for itself, and spun up a fourth agent from scratch to check the numbers. Its own estimate of the effort, in its words: "Roughly 12 to 17 hours of agent work."
Now the same six words with the page under them. Claude Fable 5.1 took the decision question straight off the page inside its first ten minutes. It planned three small data pulls inside the first half hour. It kept exactly one check for itself, and the page had named that check: whether the finance numbers in the shared drive match the August renewals. It set its own stop, and that one is short enough to quote: "At or before two hours, regardless of state, I stop." It planned one page with three options. And its estimate was "Roughly three to four hours."
Same model, same six words, same night. About four times the effort by its own numbers, and the only difference between the two runs was one page of written context.
Here is the part I did not expect, and it matters more to me than the hours. Something did not change between those two runs. In both of them the model was careful. Neither run would send anything out in the team lead's name. Neither would put a number in front of your leadership team that it could not trace. The run without the page said outright that it would not act on instructions it found sitting inside a document. So with the page underneath it, the model did less work and was no less careful.
There is a human half of this, and I constructed it rather than measured it. Hand those same six words to a capable person on your team with no page. They spend the whole week gathering anything that might matter, and your team reads them as thorough and slow. Hand that same person the page, and they come back on Wednesday with one page and three options, and your team reads them as sharp. Same person. Same ask. The only thing that moved was what somebody wrote down.
So writing those six lines is a kind of care. And the person on your team has earned more of that care than the machine, never less.
Now, the closest thing I can point you to in print comes from the companies making these models. They tell their own developers more or less what I have been telling you.
Start with OpenAI's prompting guide for GPT-5, from August 2025. It tells developers to write explicit criteria for how far the model should explore a problem. Then OpenAI introduced its next model, GPT-6 Astra, on the third of September this year. Its launch page says that model asks a focused question when the answer could change the outcome, and waits for the person when the decision is a consequential one.
And Anthropic's prompting guide for Claude Fable 5.1, the AI I use every day, hands developers a sample prompt for work the model does on its own. In that sample, the person's ask is what sets the scope.
None of that is measured. It is advice, published by the companies selling the models, with their own hedges attached. But it is the same instruction as my stop line. Say how far to go, and say when to come back and ask.
One more thing from Anthropic, and it is where I am going next. In January this year, Anthropic published a document setting out what it intends Claude to be, its values and its behavior, and Anthropic calls that document Claude's constitution. My next essay starts there.
Before I go there, the cost of writing those six lines, because you are going to be asked. The first page cost your team lead a week. The next one takes the team an afternoon. Both of them are the cheapest thing in your entire AI rollout.
I have a second line in my Cognitive Translation Protocol about what happens when nobody writes those six lines. It says, "When the artifact layer goes undesigned, the tacit assumption fills the vacuum." In plain words, when nobody has put on paper how the work gets decided, everybody works from whatever they already assume. That is exactly what your agents did with your money, and exactly what the two people from April have been doing since spring.
Now, three things make this the week to write those six lines, and I will give you each one with what it means for you.
The first. Every person on your team reads an unwritten line their own way, and The Neurodiversity Movement is the work that made that visible. Which means that page of six lines was always needed, long before a single agent showed up in your organization.
The second. You already pay for your people's thinking, and The Brain Economy is the work that attached a number to it. It calls that brain capital: brain health plus brain skills, quantified infrastructure with measurable return. Which means the operating page is how you get that return, because it is what lets the thinking you already pay for reach the work.
The third. An agent fills the same missing line with spend you can count, in dollars, on an invoice. That is what The Age of AI changed. The cost of the unwritten line was always there. Now it has a number on it, and it arrives on a Friday.
And your people are the reason any of this pays off. The people who think differently supply the judgment a model cannot. I mean the question nobody thought to ask. I mean the objection somebody put in writing before the meeting. Your organization needs that judgment sitting beside every agent it runs.
So the answer is not more AI training on its own. Duncan and Anderson wrote in Harvard Business Review in June this year that AI fluency training is necessary and far from sufficient. That is an argument from two consultants rather than a study, and I think they are right. I would say the same about your prompting workshop and your copilot certification. Keep both. Neither one is what makes your organization's judgment work.
What does make it work is design, and there is one field study I keep coming back to on that. Hackman and Wageman reported in 2005 on a field study of service teams at one company. How the team was set up accounted for 37 percent of the difference in how those teams performed. The leader's hands-on coaching accounted for under one percent. To be fair to coaching, it moved other things in that study. It just did not move what those teams delivered. That study measured how a team was set up. The step from there to writing six lines is mine, not theirs.
So your job this week is to write down how the work gets decided, for every thinker on your team, human and machine. That is the part of the design you can change without asking anyone. Then read the six lines out with as many different kinds of thinkers as the room holds.
Once that page exists, two things stop happening. You stop paying for work nobody on your team can use. And your team stops reading the thorough person as slow, because the fourth line, where to stop and check, now says how far that person should go.
And something else happens further out. When every lead in your organization has written their own six lines and a senior role opens up, the people choosing can read how each of those leads decides. That opens those roles up to people who think differently.
Your operating page names who decides. How well that person then decides is a different question, and I take that one up in my next framework, DecisionOS.
So here is the Operating Brief, the short version you can carry into your next leadership meeting. The moment, what your organization does today, the redesign, and the test for Monday.
The moment. Your pilot's agents are running. The spend is climbing. You cannot show one result. Those agents never had an operating page, and neither did the team running them.
What your organization does today, without anyone having chosen it. That page of six lines stays unwritten. So the agents keep filling the missing lines with spend, and your people keep filling the same lines with guesswork. Your people guess at what is locked and who decides, and you are paying them for their judgment. And the answer your organization reaches for is more seats and more training. The prompting workshop. The copilot certification.
The redesign. Every thinker in your organization, human or machine, runs on what somebody wrote down. So you write that clarity once, for all of them, on one operating page of six lines. What matters, and what is not being decided. What is known, and what is assumed. What is locked. Where to stop and check. What to drop if time runs short. And who decides.
And the Monday test. Take the ask you would hand an agent this week, and write its six lines in plain words. Then hand that same page to the people on your team. And you need nobody's permission to do any of it.
One more thing, and it takes two minutes. There is a prompt printed in the essay. Paste it into Claude, or whatever AI your organization uses. Then paste your six lines underneath it, or just the bare ask if you have not written them yet. Strip the names out first.
That prompt asks the AI to read your page the way a capable stranger would, someone new to your team with nothing else to go on. What comes back is what it still cannot know, as plain questions you have to answer before work starts. That is the list your team has been guessing at.
The full argument is in the essay on the site. It carries both of those AI replies printed word for word, and every source I have named here with what it does and does not show. My framework page has the whole Cognitive Translation Protocol on it.
If this gave you something to run on Monday, subscribe to the Substack. Each week's essay lands there first, and it is the one place I ask you to follow.
Your team lead's first page cost a week. Yours will cost an afternoon. Every thinker in your organization, human or machine, runs on what somebody wrote down. The first page is yours to write this week.