Two official signs at the same mountain pass, four hundred feet apart, both wrong. Nobody retires anything — including me, and I wrote the protocol.
There are two signs at the top of the Khardung La. They stand a few metres apart. They disagree with each other by four hundred feet.
The Border Roads milestone says 17,982 feet. The green health advisory board beside it says 18,380 feet / 5,602 metres. Every GPS survey, every satellite elevation dataset and every serious piece of mapping puts the pass at 17,582 — about the height of the Everest base camp on the Nepal side, a couple of hundred metres above the one in Tibet.
Read that again, because it is the whole article.
They did not correct the sign. They left the old one standing, made a new one, and still got it wrong. The newer board also warns you about “low oxygne levels.” Nobody proofread that either.
Both are official. Both are freshly painted. Both are photographed a thousand times a season by people who ride up, take the picture, buy the T-shirt, and carry the number home.
Now ask yourself how often that happens where you work.
Nobody retires anything. We layer. The new policy goes on top of the old policy and the old one never comes down.
The runbook gets revised and the induction pack does not. The checklist survives three reorganisations and two systems. Two controlled documents sit in force, contradicting each other, and nobody notices — because each one looks perfectly authoritative on its own.
And the new version, the one written specifically to fix the problem, never gets audited at all. It is new. New things are assumed correct. That assumption is the single most expensive one in this article.
What I said in April.
In April I published something uncomfortable. I asked an AI to assess a twelve-hour build — thirteen merged pull requests, four waves, zero rollbacks — and then to score my career against the standard that build had just set.
It gave me 10 out of 100.
That was not a humblebrag and it was not false modesty. It was a measurement, and I published it because the number was useful. The conclusion I drew then still stands:
The AI was the multiplier. The protocol was the multiplier on the multiplier. And the protocol is human.
I wrote that line. I believed it. I put it in front of eleven thousand people.
Four months later I walked straight into it.
What I did in August.
Five platforms. More than a million lines of code between them. Pre-production launch cycles running in parallel, each burning around 1.5 million tokens a pass.
And they worked. Every run made solid ground. Each platform came within a whisker of launch certification — close enough that you can taste it, close enough that you start drafting the announcement.
Then every one of them hit a wall.
Not a clean board. Not one run where every test came back green. And here is the detail I should have caught on day one and did not catch for three weeks: it was the same wall every time.
Random failure is a capability problem. Repeatable failure at the same point is a specification problem. The machine was not struggling. It was doing precisely what it had been told, precisely as well as it had been told to, and arriving at the same place — because the instruction set had a fault in it and the instruction set never changed.
I was reviewing output daily. I never once reviewed the instrument doing the measuring.
The embarrassing part.
I have spent twenty-five years in and around C-suites and boardrooms refusing to accept because the policy says so as an answer.
Every board I have sat on, every executive team I have run, would tell you the same thing about me: I do not take a governance document as the end of a conversation. I have rewritten constitutions, torn up pay frameworks, dissolved structures that everyone agreed were permanent, and authored a national reform that dismantled the very system I was operating inside. Not settling for the status quo because it was written down somewhere is, more than anything else, the through-line of my career.
Then I wrote a protocol of my own. And for three weeks I treated it as scripture — the newest document in the building, and therefore, obviously, the right one.
The V1 runbook sat on the table for three weeks. That is not a bad afternoon. That is three weeks of runs, ground I never took, and a token bill Anthropic is delighted with and I am not.
I spent a career refusing other people’s rules and three weeks obeying my own.
That is the whole failure. Not compute. Not context windows. Not model choice. I had built a carefully calibrated human protocol, it had worked brilliantly, and somewhere between the success in April and the wall in August it quietly became the thing I had spent my whole career warning other people about: a document that had stopped being questioned because it used to be right.
I named the trap myself. Legacy thinking. Then I fell in it.
Why this is not a story about me.
There is a rush of reporting at the moment about companies that are not seeing the returns on their AI spend and are pulling back. Boards are asking harder questions. Budgets are being cut. The mood has turned.
The most cited piece of evidence is MIT’s The GenAI Divide: State of AI in Business 2025 — the study that found roughly 95% of enterprise generative-AI pilots producing no measurable return against $30–40 billion of investment.
Read past the headline and the finding is not what most people took from it. The report does not blame the models. It blames a learning gap — systems that, in its words, do not retain feedback, adapt to context, or improve over time.
That is exactly right, and it is still one level too shallow.
The loop that is failing to learn is not the machine’s. It is the one wrapped around it. The protocol. The escalation path. The definition of done. The quality gate. The thing a human wrote, six weeks ago, when they knew less than they know now — and has not opened since.
Organisations have decades of muscle for auditing output. Variance analysis. QA. Post-implementation review. Internal audit. We are extremely good at asking did the thing we built work?
We have almost no muscle for auditing the instruction set, at anything like the cadence this era demands. And output audit will never find this class of failure, because the output is not wrong. The output is a perfect execution of a flawed specification, delivered on time, every time.
The machine delivers against your calibrated schema. It will deliver the wrong thing flawlessly, and it will never tell you.
Companies cutting AI budgets because the returns are not there are drawing the same conclusion I drew on that pass: the climb is not working. It is the correct observation and the wrong diagnosis. They are stuck in the loop, and the loop is human.
What changed.
I came down. I stopped looking at the output and opened the protocol itself — and I did the one thing I had not done: I put a different intelligence on it. I rewrote the QC protocols with Opus and re-orchestrated the production release cycle from the schema up, rather than patching what kept breaking.
Same five platforms. Same roughly 1.5 million tokens. One simultaneous run.
A week of ground, in about an hour.
Nothing changed in the model. Nothing changed in the compute, the budget or the team. The only thing that changed was the quality of the instruction it was executing against — and that was mine to fix the entire time.
The half-life.
Here is the part I would put in front of a board.
Every operating protocol has a half-life. That is not new. What is new is the number.
An operating model used to last long enough to be written down, socialised, put in an induction pack and reviewed at the next planning cycle. Years. Long enough that annual review was a reasonable cadence and governance could sensibly mean stability.
In this environment the half-life of a protocol is measured in weeks. Mine had a good run and then went out of date while I was on a mountain — and it went out of date not because it was badly written but because the ground under it moved, which is now the normal condition rather than the exception.
Four things follow from that, and none of them are technology decisions:
- Put an expiry date on the protocol, not just the plan. If your AI operating protocol has no review date, it is already expired. Treat it like a controlled document with a short shelf life, because that is what it is.
- Separate the two audits. Auditing output tells you whether the machine did what you asked. Auditing the protocol tells you whether you asked for the right thing. Most organisations only run the first. Only the second finds this failure.
- Make the author the least trusted reviewer. Nobody catches their own specification error — I certainly did not, and I wrote the sentence warning about it. The protocol review belongs to someone who did not write it. A different person, a different model, an outside pair of eyes, but not the author.
- Measure recalibration time. How long between this stopped working and the protocol changed? That interval is now a real operating metric, and in most organisations nobody owns it, nobody measures it, and it is counted in quarters.
The rate.
None of this is really about AI.
It is about the rate at which human beings are being asked to update their own operating assumptions — and that rate is now higher than at any point in our history. The industrial revolution gave people a generation to adapt. Electrification, decades. The PC, a career. The internet, the better part of one.
This asks for a rewrite every few weeks, and it asks it of the same nervous system that evolved to find a rule, learn it, trust it, and stop spending energy on it. Writing the rule down and then not thinking about it again is not a character flaw. It is the single most efficient thing the human brain does. It is also, right now, the most expensive.
The protocol is human. I still believe that — more than I did in April.
But I left half the sentence out. The protocol is human, which means it is fallible, it is dated the moment it is written, and it is revisable. The first two happen to you. The third is a decision, and it is the only part of this you actually control.
There are still two signs at the top of the Khardung La, four hundred feet apart, and a third number — the true one — on nobody’s board at all. One of them tells you not to stay more than five minutes. It has been wrong for thirty years. I have a photograph of myself grinning next to each of them.
Check your own signs. Especially the ones you put up to replace the last set.
For my kids and yours.
Series: Symphony (April 2026) → Vibe Coding, Vibe Engineering, The New Leadership Archetype → The Pre-Orchestration Era.