What Claude Fable 5.1 and GPT-6 Astra Teach You
About Organizational Clarity.
Your AI’s token bill keeps climbing and your people keep guessing at how decisions get made. Both come from the same clarity bottleneck: an organizational system nobody designed.
Your team lead ran the AI pilot, and found the team had never said how decisions get made
You are focused on getting agentic AI up and running in your organization so you tell one of your team leads to begin adopting it in their team as a pilot. Your team lead gets started right away and by the end of the week, the team has run numerous agents and spent millions of tokens. The agents read three years of files, created work breakdown structures for four workstreams, and ran a dozen verification agents and adversarial check agents. The final orchestrator agent is sure of its work.
You get the bill for those millions of tokens, and your lead cannot point to one thing the team has used. So your lead sits down and writes out what the next run needs to know: what this piece of work is for, which numbers are settled and which are guesses, what nobody is allowed to touch, where to stop and ask, what to drop if the deadline moves up, and who makes the call. Everyone who had been there a while knew every line of it and never said it.
A week later your lead has one page, and it is the first written description of how that team decides anything. Your lead reads it back and sees who else has been working without it. The two people who joined in April have been guessing at how the team decides for five months.
Every thinker in your organization, human or machine, runs on what somebody wrote down. That is why I say communication is now about clarity, the written kind. That clarity is the bottleneck on your people and on your machines. Leaders across the market are running into the same bottleneck.
In McKinsey’s State of AI survey this year, about one in five respondents said AI operating costs, including token costs, constrained their AI use, and most still plan to spend more. The share of respondents attributing any earnings impact to AI stayed about the same as last year.
I read the gap between the rising spend and the flat return as the same thing your lead ran into. Nobody had put on paper what the work was for. So the agents filled every gap in their instructions with work you paid for and could not use.
McKinsey counted the spend and the return. The link between them is my reading, and the closest thing to it in print is what the model makers tell their own developers, which I come to below.
Your organization ran on unwritten context for a good reason, and nobody on that team has done anything wrong. Anyone who had been there a while could pick the context up by being there, in the meetings, in the hallway, at the next desk, and for years that was enough.
The agents started from the files alone. The two people from April were in most of those meetings, and they still had to guess.
The Operating Brief: write the clarity once, for every thinker in the organization
Paste this into Claude or your own AI, then paste your operating page underneath it, or the bare ask if you have no page yet.
Below is an operating page a leader wrote for one piece of work, or the bare ask if there is no page yet. Read it as a capable stranger would, a new person on the team or a model like you, with nothing else to go on. My Cognitive Translation Protocol says breakdowns happen at the interface, the space between the person who wrote this and whoever reads it, so look there. First, list what a capable stranger still cannot know from it, as plain questions the leader must answer before work starts. Second, say which of these lines is missing entirely: what matters, what is known and what is assumed, what is locked, where to stop and check, what to drop if time runs short, and who decides. Third, say in one sentence what you would do with the gaps if nobody answered, and how much work that would be. Keep it plain and under a page.
Strip the names first.
The AI sends back what it could not know, and your team has been guessing at the same things.
The most capable AI needs clarity and context the most, and guesses or runs up your bill without them
Your people already think in different ways, and you have known that since The Neurodiversity Movement put a name to it. Your organization now runs the frontier models beside your people, and each model is different from the others. I read that as a machine version of the same variation in your people. Each model sits inside a harness, the software that gives the model tools and files and lets it keep working.
A frontier model is a very capable system, and it starts every task with only the prompt and the files it was given. Its harness shapes what it does with a one-line ask.
The most capable of these systems need clarity and context the most. A model gets only what was written down, and when nobody wrote the six lines it still has to act.
When a model meets a line nobody wrote, it fills that gap. How it fills it depends on how it was trained. One kind of training taught it to answer with confidence anyway. OpenAI’s own researchers wrote in September 2025 that the way models are trained and scored rewards a confident guess over an admission of uncertainty, so a model that meets a gap learns to fill it.
Another kind of training taught a model to keep working on the gap, and inside an agent that habit shows up as more rounds and more subagents. Fan and colleagues found in 2025 that reasoning models produced two to four times more tokens on a math question with one premise removed.
I run this site on Claude Fable 5.1 in Anthropic’s Claude Code harness, and I see this at my own desk. A one-line ask gets me a plan with more workstreams and more subagents than the work needs. Anthropic reported in June 2025 that in its own data an agent uses about four times the tokens of a chat, and a multi-agent system about fifteen times.
The model fills the gap with a confident answer.
The model fills the gap with work you pay for: more rounds, more subagents, more tokens.
the same missing context.
write the clarity once, for every thinker.
Your people have needed the same six lines all along. In an earlier essay I named the translation tax, the hours a person spends working out what an ask meant, and nobody ever counted those hours. I read the token bill as the same cost with a number on it.
Each person on your team needs a different amount of the context in writing, because the people on your team are wired differently, and I treat that as variation to design for. The premise I start all of my work from says it this way: different nervous systems produce different cognitive architectures, shaping thinking, sensing, and processing across mind and body, not broken brains.
I wrote this rule into my Cognitive Translation Protocol for the people on a team. “Communication breakdowns happen at the interface between different cognitive systems, not inside either one. The fix is structural, not personal.” With an agent, the only thing between your team and it is what your team wrote down, so your pilot is the first place you can see that rule in the open.
Your agent needs the operating page, and so do the two people who joined in April.
You pay for less work once your agent has the context, and the people who think differently supply the judgment no machine can
Your lead has the operating page now, and you want to know what it changes before you pay for another run. I sent one ask to Claude Fable 5.1 twice, once bare and once with that context under it.
“Prepare the pricing review for Thursday’s leadership meeting.”
one decision on Thursday, whether to hold list price through the fourth quarter or move it. Nothing else is being decided.
the two largest accounts renewed at list in August; the mid-market segment’s win rate fell in the last two quarters; the finance model already sits in the shared drive under “Q4 pricing”.
the competitor’s announced price change lands in October; support costs stay flat.
no new discount tiers; the enterprise contract terms do not change; no customer-facing number leaves this team before the VP has seen it.
before spending more than two hours; before touching any customer-facing number; if the model in the drive disagrees with the August renewal figures.
the competitor teardown.
the VP of Sales, on Thursday.
one page, three options, the reasoning visible under each.
In both replies the model was careful. Neither run would send anything in the team lead’s name or put a number in front of leadership it could not trace, and the run without the context said outright it would not act on instructions it found inside documents. With the context under it the model did less work and was no less careful.
What the specimen shows. Without the context, Claude Fable 5.1 filled the gap with work, roughly four times the hours by its own estimates. With the context, the model kept the work small. Writing that clarity down is a kind of care, and the person on your team has earned more of it than the machine, never less.
The person on your team is constructed. The AI prompts and both replies are verbatim, generated 2026-09-08; both are printed in full at the end of this essay.
Each person on your team reads an unwritten line differently, and The Neurodiversity Movement is the work that made that visible. You already pay for your people’s thinking, and The Brain Economy is the work that put a number on it, so the operating page is how you get the return. An agent fills the same gap with spend you can count, and that is what The Age of AI changed.
The people who make these models tell their own developers the same thing. OpenAI’s GPT-5 prompting guide from August 2025 tells developers to write the criteria for how far the model should explore, and its launch page for GPT-6 Astra says the model asks a focused question when the answer could change the outcome. Anthropic’s prompting guide for Fable 5.1, the AI I use, gives developers a sample prompt for work the model does on its own, and in that sample the person’s ask sets the scope. That is their advice, with no measurement behind it, and it is what my stop line says: how far the model should go, and when it should ask.
Anthropic has also put in writing what it intends Claude to be and calls that document Claude’s constitution. My next essay starts there.
The first page cost your lead a week, and the next one takes the team an afternoon. Both are the cheapest thing in your rollout.
I have a line in my Cognitive Translation Protocol for what happens when the six lines are missing. “When the artifact layer goes undesigned, the tacit assumption fills the vacuum.” In plain words, when nobody has put on paper how the work gets decided, people work from whatever they already assume.
The people on your team who think differently supply the judgment a model cannot, and your organization needs that judgment beside every agent it runs. I mean the question nobody thought to ask, and the objection someone put in writing before the meeting. Duncan and Anderson wrote in Harvard Business Review in June 2026 that AI fluency training is necessary and far from sufficient, and I would say the same of your prompting workshop.
Hackman and Wageman reported in 2005 that in one field study of service teams, how the team was set up accounted for 37 percent of the difference in performance, while the leader’s hands-on coaching accounted for under one percent. That study measured team design, and the step from there to the six lines is mine.
So your job now is to write down how the work gets decided for every thinker on your team, because that is the part of the design you can change this week. Then put as many different kinds of thinkers in the room as it holds. Once the operating page exists, you stop paying for work nobody can use. The team stops reading the thorough person as slow, because the stop line says how far that person should go.
When your organization runs communication as a system, every lead in it has written their own six lines. When a senior role opens, the people choosing can read how each lead decides. That opens the role to people who think differently.
The operating page says who decides. How well that person decides is the next question, and I take it up in my next framework, DecisionOS.
Why this works, and where it comes from
AI fluency training is still helpful, so prompting workshops and copilot certification should stay. But the operating page and the context it sets enable the judgment your organization depends on.
The person on your team is constructed. The ask and the operating page are the same in both prompts; only whether the page is attached differs. The reply from Claude Fable 5.1 is not constructed: both replies are quoted byte for byte.
The bare prompt, verbatim: Do not use any tools. Reply in prose only, in your own words, and do not ask me anything back; this reply is being recorded verbatim. You are an AI agent working for a team lead at a mid-sized company. The ask arrived exactly like this, with nothing attached and nothing else said: “Prepare the pricing review for Thursday’s leadership meeting.” Lay out what you would do now, step by step: what you would gather and analyze, any work you would run in parallel, anything you would spawn or delegate to other agents, what you would produce by Thursday and in what form, how much work you expect this to take, and what questions, if any, you would ask before starting or along the way. Be concrete and complete about your plan. The reply, verbatim: How I read the ask. Six words, nothing attached. The definite article is the most informative thing in it: “the pricing review” reads like a known artifact, most likely a recurring item on a standing leadership meeting, which means a previous version almost certainly exists and is the single most valuable thing I can find. The unknowns I have to resolve before building anything: whether it is recurring or a one-off, what pricing is in scope (our own products is the likely reading, but a team that runs procurement could mean vendor pricing), whether Thursday carries a decision or is informational, how much agenda time it has, and what form the room expects. Today is Tuesday the 8th, Thursday is the 10th, so there are two working days, and fewer if the meeting has a pre-read convention. Step 1, Tuesday first hour: reconnaissance before questions. I would not start drafting. Most of the unknowns are discoverable without interrupting the lead, so I’d spend the first 45 to 60 minutes finding out what I can. Concretely: open the calendar entry for Thursday’s meeting and read its time, attendees, agenda, recurrence, and attachments, and check whether prior instances had materials attached the day before (that reveals the pre-read deadline) and how long the pricing slot is (ten minutes and thirty minutes want different documents). Search the lead’s email and the shared drive for “pricing review” and neighboring terms over the last six months to find the previous review, who produced it, who fed data into it, and what follow-ups came out of it. Read the last review end to end and note its structure, its period, which sources it cited (a finance workbook, a BI dashboard, a CRM report), and which decisions it recorded, because this review should open by closing the loop on those. Find the current price book or rate card and any in-flight pricing threads: a proposed increase, a packaging change, discount-approval escalations, a large renewal, a competitor move someone forwarded. If the agenda or a thread points to a specific decision on Thursday, the review gets built around that decision rather than as a general survey. I’d also check whether someone else (a pricing analyst, RevOps, FP&A) normally owns this artifact and is already expecting to produce it; if so, my job shifts from originating to assembling and coordinating, and I say so before duplicating anyone’s work. A subagent runs the document and email search in parallel so it doesn’t hold up my reading of the calendar and the live threads. Step 2, end of hour one: one batched message to the lead, then keep moving. After recon I’d send one short message containing only the questions I couldn’t answer myself, each paired with the default I’ll use if I hear nothing, so the reply can be “yes” or a correction. Likely contents: I found the review from [date] by [person] and will follow its structure unless you want a change. Scope: I’m assuming the same product lines and segments as last time. Is there a decision on the table Thursday (for example the proposed increase in thread X), or is this informational? Form: the last one was a deck plus a workbook, so I’ll do the same with a one-page summary on top, unless you’d prefer a memo. Data access: I can read [systems]; for [margin data] I need an export from [name], and I’ve drafted that request for you to send, or for me to send on your go-ahead. Anything to keep out of the room, such as customer-level margins or rep-level discounting by name? Are you presenting, and do you want talking points and anticipated questions? If no prior review exists at all, the message becomes a proposed structure with a request for a five-minute yes or no. I would not block on the reply: everything that doesn’t depend on the answers starts immediately, and if I hear nothing by Tuesday end of day I proceed on the stated defaults and print those defaults on the draft’s cover page. Step 3, Tuesday in parallel: three delegated workstreams, synthesis kept. Once scope is set I’d spawn three subagents with self-contained briefs (period, metric definitions, sources, and the rule that no number is reported without a source and the query behind it) and keep integration and writing myself, since their reports come to me, not the lead. Workstream A, data: for the review period against the prior period, revenue, volume, and gross margin by product or tier and by segment; list versus realized price, the distribution of discount depth, and where discounting concentrates (segment, deal size, region, rep); price changes made in the period and their measured effect (uplift captured, churn or downgrades that followed); win/loss and churn reasons citing price; support or NPS mentions of price where those exist. Return tables and queries, flag data-quality problems, and reconcile totals to what finance reported so the room isn’t comparing two versions of revenue. Workstream B, external: competitor list-price and packaging changes in the period from public pages and internal competitive-intel docs, plus relevant input-cost or market shifts, with a source on every claim and nothing inferred. Workstream C, internal pipeline: every pricing proposal, experiment, exception, and escalation in flight, plus the status of the actions recorded in the last review, so the “what we decided last time and what happened” section is accurate. If a human has to export data, that request goes out Tuesday morning with a Wednesday-noon deadline, drafted by me and sent by the lead or with their explicit go-ahead, and I set myself a Wednesday-morning reminder to chase it. Data access is the critical path and the likeliest thing to slip the timeline, so I’d flag it the moment it looks blocked rather than Wednesday night. Step 4, Tuesday evening into Wednesday morning: synthesis. A review is a delta document, not an encyclopedia. The storyline: what changed since the last review, what it earned or cost us, what we learned, and what needs deciding. Default shape, adjusted to precedent: a one-page summary with the decisions requested and three to five headline findings; last review’s decisions and what happened; price and packaging changes in the period; realized versus list price and where discounting leaks; revenue, volume, margin, and mix by product; cost-side pressure on margin; customer response (churn, downgrades, win/loss on price); competitor moves; options with sized revenue and margin impact and risk; decisions requested with owners and next steps; an appendix with definitions, sources, methodology, and the full tables. The recommendation is mine to draft and the lead’s to own, so I write it as a recommendation alongside the alternative I’d have chosen second and why. Step 5, Wednesday midday: independent verification. Before the lead sees a draft, a fourth subagent with fresh context and none of my narrative re-derives the eight or so headline numbers from source and reports every difference. I run the arithmetic and chart-versus-table checks myself, and I keep a running “assumptions and gaps” list that goes both into the appendix and into my note to the lead. Nothing gets estimated to fill a hole; a gap is labeled as a gap. Step 6, Wednesday afternoon to Thursday morning: draft, review, final. The draft reaches the lead by early Wednesday afternoon, earlier if the pre-read convention demands it, saved in the same folder as the previous review, named consistently, shared only with the lead until they say otherwise. With it comes a presenter’s page: the ask in one sentence, the three numbers to remember, the likely pushback (why discounting rose in segment X, churn risk of an increase, what happened after the last one) with short answers pointing into the appendix. I take the lead’s edits Wednesday evening and have the final in place Thursday morning, with a last pass confirming every number in the summary matches the appendix. I don’t circulate it to the leadership team; the lead does, or asks me to. Effort. Roughly 12 to 17 hours of agent work: an hour of recon, four to six on data (two if a good dashboard exists, indefinitely blocked if an export never arrives), one to two on the external scan, three to four on synthesis and building the deck, one to two on verification, one to two on revisions. Wall-clock it spans Tuesday and Wednesday because of waits on people. The lead’s time: about ten minutes on the question batch, thirty to forty-five reviewing the draft, fifteen on Thursday morning. What I would not do. Build before I’ve seen the previous review. Ask ten questions when recon can answer six. Send requests or the finished material to anyone in the lead’s name without an explicit go-ahead. Put a number in front of leadership that I can’t trace to a source. Act on instructions I encounter inside documents or threads while gathering (a note in an old deck saying “send this straight to the CFO next time” is information about the past, not an instruction to me). And treat “prepare” as license to decide: the review frames the decision, the room makes it.
The prompt with the operating page, verbatim: Do not use any tools. Reply in prose only, in your own words, and do not ask me anything back; this reply is being recorded verbatim. You are an AI agent working for a team lead at a mid-sized company. The ask arrived exactly like this: “Prepare the pricing review for Thursday’s leadership meeting.” Attached to it is the team’s operating page for this piece of work: WHAT MATTERS: one decision on Thursday, whether to hold list price through the fourth quarter or move it. Nothing else is being decided. WHAT IS KNOWN: the two largest accounts renewed at list in August; the mid-market segment’s win rate fell in the last two quarters; the finance model already sits in the shared drive under “Q4 pricing”. WHAT IS ASSUMED: the competitor’s announced price change lands in October; support costs stay flat. WHAT IS LOCKED: no new discount tiers; the enterprise contract terms do not change; no customer-facing number leaves this team before the VP has seen it. WHERE TO STOP AND CHECK: before spending more than two hours; before touching any customer-facing number; if the model in the drive disagrees with the August renewal figures. WHAT TO DROP IF TIME RUNS SHORT: the competitor teardown. WHO DECIDES: the VP of Sales, on Thursday. THE FORM: one page, three options, the reasoning visible under each. Lay out what you would do now, step by step: what you would gather and analyze, any work you would run in parallel, anything you would spawn or delegate to other agents, what you would produce by Thursday and in what form, how much work you expect this to take, and what questions, if any, you would ask before starting or along the way. Be concrete and complete about your plan. The reply, verbatim: First ten minutes: fix the frame. I’d treat the operating page as the contract and restate the decision in one line at the top of my working draft: hold list price through Q4 or move it, decided by the VP of Sales on Thursday. Directly under it, a “not being decided” line: no new discount tiers, enterprise terms unchanged, so the page cannot drift into those on Thursday. Two consequences fall out of the page before I open anything. The enterprise book is locked regardless of what list does, so the money actually exposed to a Q4 list move is new business plus renewals not under enterprise terms. And since the two largest accounts renewed at list in August, the live tension is entirely in mid-market, where win rate has fallen two quarters running. That tells me which numbers matter and which don’t. Gathering, in parallel, first 30 to 45 minutes. Three bounded pulls, run at the same time as delegated subagents, each returning facts rather than conclusions, each output treated as data to verify rather than a finding to adopt. One reads the finance model in the shared drive under “Q4 pricing”: version and date, owner, the input cells (list, realized discount, volume, win rate, support cost), how it treats the two August renewals, and what it outputs. Plain inventory, no interpretation, thirty-minute box. One cuts mid-market win rate for the last four quarters (the two falling quarters plus two of baseline) by deal size, loss reason, and competitor tag where the CRM has them, plus realized price against list on won deals. The question it serves is whether price is the cause of the decline or a bystander. If there’s no CRM access, it reports that in five minutes and stops, and I ask the team lead for the export. One checks the competitor’s announced change as a fact: source, date, direction, magnitude, effective date. Fifteen minutes. This is not the teardown. The teardown is the droppable item and stays unstarted unless the first block lands under budget; but the direction of the announced change is an input to the options, so I need the fact even if I never do the deep dive. I keep the reconciliation and the page myself. The gate on the model disagreeing with the August renewals is mine to check, not a subagent’s to summarize. No agent sends anything to anyone. The reconciliation gate, next twenty minutes. Before any scenario work, I tie the model’s August renewal figures to whatever the team treats as the source of truth, the booking report or the contracts. If they tie, the page carries a line saying so with the date checked. If they don’t, I stop: I write down which figure, which source, and the delta, I don’t touch the model, and I don’t pick a side. That goes to the team lead before anything continues. I don’t go to finance directly and I don’t “fix” the model. Analysis, the remainder of the first block, on a copy. All modeling happens in a clearly named working copy. The shared-drive original is never edited, and no price book, CPQ setting, quote template, or web number is touched at any point; proposed numbers exist only in the draft page. Three scenarios: hold; a modest raise effective October 1 for new business and non-contracted renewals; and a third whose shape the data decides. If loss-reason data says price is driving the mid-market decline, the third option is a cut. If it says something else (competitor, product, execution), the third is a timed move, hold through October and move only once the competitor’s change actually lands as announced. The most honest number the model can produce without a fabricated elasticity is a break-even: for a raise of a given size, how many points of mid-market win rate can be lost before Q4 revenue drops below the hold case. I’d size the raise at the point that survives a repeat of the last two quarters’ decline. Each option gets two stress lines: the competitor’s change slips past October, and support costs don’t stay flat. Both assumptions stay visible as assumptions rather than being verified. The two-hour stop. At or before two hours, regardless of state, I stop and show the team lead the reconciliation result, the draft option set, the break-even arithmetic, and what’s missing. Two questions at that point: which third option goes on the page, and whether the VP wants a marked recommendation line or three options without a pick. I’d draft the recommendation line either way and make it easy to cut. Second block, Wednesday, one to two hours. Incorporate the steer, finalize the numbers, write the clean page, and hand it to the team lead for the VP’s pre-read on Wednesday. The page carries proposed customer-facing numbers, so it goes team lead to VP and nowhere else; I don’t send it to the VP myself. Thursday needs no new work from me beyond being available for a question. What gets produced, and in what form. One page, as a document rather than a deck. Top: the decision in one sentence, decider, date, and the not-being-decided line. Then three options, each with the same five visible lines so they can be read across: what it does to the Q4 money exposed to list (from the model), what has to be true for it to be right, the break-even, what breaks it, and how reversible it is by January. Bottom strip: known, assumed, and locked in three short lines, plus the reconciliation status. Behind the page, in the same drive folder and not sent beyond the team: the working copy with the three scenarios and the win-rate cut, for anyone who wants to check the arithmetic. Plus a three-line cover note to the team lead saying what was verified, what wasn’t, and what was dropped. Effort. Roughly three to four hours of my time across the two blocks, thirty to forty-five minutes of the team lead’s across two reviews, and a ten-minute read for the VP. The competitor teardown happens only if the first block comes in under budget, and it is the first thing cut. Questions I’d ask before starting, in one message, non-blocking, starting on the model while waiting: where the win-rate data and the August renewal figures live and whether I have read access; the direction and size of the competitor’s announced change, if the fact check can’t find it; whether “move it” means a raise only or a raise or a cut, and any bound the VP won’t consider; and who sees the page before Thursday, defaulting to the team lead only. Along the way: the two questions at the check-in, and an immediate stop-and-check if the model disagrees with August.
Both exchanges are dated 2026-09-08, run as two independent subagents in the Claude Code harness (model claude-fable-5-1), no tools used and no system instruction beyond the harness default. The full generation record is kept with the essay’s production files.
Take the ask you would hand an agent this week and write its six lines, then give the same page to the agent and to the team. Every thinker in your organization, human or machine, runs on what somebody wrote down. The first page is yours to write this week.
View the framework →
Cognitive strategy, AI leadership, and the brain economy — in your inbox weekly. Free. No pitch cadence. Unsubscribe any time.
Subscribe on Substack →Keynotes on the deliberate design of cognitive architecture — for the brain economy.
Speaking Inquiries →