There is a reasonable objection to the whole argument about rethinking processes and roles, and it deserves to be taken seriously, because most of the time it is right. Whoever proposes to revise the process is almost always proposing to slow down. They add an approval step, a document to fill in, a committee that meets on Thursdays, and they do it at the worst possible moment, when a team has finally found a rhythm. Every methodology of the last thirty years arrived as quality and was delivered as bureaucracy. The AI governance pages produced over the last two years are largely theatre, and they exist so somebody can tell a board the topic is covered. If someone shows up holding a reorganisation proposal, the suspicion is healthy.
A Failure That Throws No Exception
What the objection does not cover is a class of failures that produces no errors. It raises no ticket, fails no pipeline, appears on no dashboard. The symptom is not shaped like an incident but like a sentence said in a retrospective, usually around the third or fourth month of working with an agentic tool, when somebody observes that the model has got worse.
The natural reading is that the vendor degraded something, or that the early enthusiasm was inflated. Almost always that is not what happened. What happened is that the project’s instruction file grew past the point where the agent can hold all of it, and started quietly dropping the half that mattered. Or that the description of an automated procedure drifted through rewording after rewording until it fires on anything, which amounts to firing on nothing. Neither failure throws an exception. There is no stack trace for scaffolding coming apart, and with no error to look at, the team attributes the drop to the tool instead of to its own scaffolding.
The Status of Personal Configuration
This is the first appearance of a principle that, once seen, shows up in three or four different places inside the same company. What degrades quietly is exactly what nobody is required to check.
Application code has a defence against silent decay, and the defence is that somebody wrote a test that gets angry. The instructions we steer an agent with, the saved prompts, the project conventions, the automations that are supposed to fire on their own, all that material which over the last two years became production infrastructure, inherited the status of personal configuration. It lives in a project folder with no declared owner, no changelog, nobody answering for its maintenance. In practice, the youngest and most fragile scaffolding in the whole delivery chain is also the only one we never thought to put under verification.
The Word “Done”
The second appearance concerns a word we treat as defined and is not. The word is “done”. A year and a half of generative tooling has accelerated a great many things, and one of the things it has accelerated most is the completion claim. An agent reaches the end of a task with a confidence that has no counterpart in verification, and the sentence it closes with is indistinguishable from the one you get after work that was actually finished.
If the definition of done is written in prose, in a page of internal conventions everybody read once, it gets honoured when there is time and skipped when there isn’t, which is when it matters. It is not skipped out of negligence. It is skipped because a rule in prose will always compete, and always lose, against a delivery that has a date. The difference between a written rule and a deterministic check that blocks the completion claim when a source file changed after the last passing test is not a difference of rigour. It is the difference between an intention and a constraint, and under pressure only the second survives.
Meanwhile the cost has not vanished, it has moved. It moved to review, to rework, to the reopening of tasks already marked closed, and in none of those three lines is anybody measuring it, because the metric we keep watching is the speed at which things get declared finished.
The Inventory Nobody Regenerates Mid-Sprint
The third appearance is the one I know best, because I have spent years working with clients in healthcare and public administration, where compliance is not a reputational topic but a condition of admissibility. Regulatory compliance does not fail through ignorance. Nobody, in any company that has read a public tender in the last three years, doubts that a software bill of materials is needed. The point is that nobody regenerates it mid-sprint, because nothing forces that moment to exist: there is a release to ship, and regenerating the dependency inventory competes with that release for somebody’s attention and loses, every time, until the day an auditor asks for it and it isn’t there. I made the case at length arguing that compliance does not fail for lack of rules, and I have not found a single counterexample since.
The failure is identical to the previous two. It is not a competence gap, it is the absence of a forced moment. The difference is that here the clock is public and not negotiable, because Cyber Resilience Act vulnerability-reporting obligations apply from 11 September 2026, the regulation applies in full from 11 December 2027, and the European Accessibility Act has been in force since 28 June 2025. Friction that until yesterday cost a rhetorical figure in a meeting now costs a non-conformity.
Buying Capacity Instead of Diagnosing
A fourth occurrence is worth adding, smaller but diagnostic, because it shows how easily the cure is mistaken. When a team starts hitting the usage limits of agentic tools, the immediate conclusion is that a bigger plan or a different model is required. In the large majority of cases the cost driver is not the model chosen, it is the context accumulated over a long session, and the fix is not in the purchase but in how the work is split.
It is the same error at reduced scale: faced with an invisible failure, we buy capacity instead of diagnosing the mechanism, because buying capacity is something you can decide in ten minutes and diagnosing a mechanism is not.
Who Can Say No
From here the question stops being technical and becomes a question of roles, which is why I find it interesting. If the bottleneck moves from production to verification, the scarce resource is no longer whoever writes fast. It is whoever can say no with a reason, and can say it about an artifact that has every appearance of being correct. It is the same migration of value I have called specification debt, seen from the output side rather than the input side.
In small teams this produces an effect I have watched from close up: a junior working alongside an agent delivers output with the surface of senior work, without the judgement that usually comes with that surface, and the review load concentrates on two or three people who become the constraint on everything else. The instinctive reaction is to ask those two or three people to review faster. The right reaction is to ask which part of their judgement was in fact a rule nobody had ever written down, and to write it in a form that fires on its own.
Training and the Floor Are Two Investments
This is where training and baseline separate, and they need funding as two distinct investments because they answer different questions. Training raises people, widens what they can see, moves somebody from the level of execution to the level of specification and verification. But training produces results above a floor, and in a team of ten the floor does not exist by default. It exists only if somebody builds it, puts it under version control and treats it as a maintained artifact rather than a folder of notes.
Without that floor below, raising the people stays a line in an institutional slide deck: everyone works at their own level of rigour, the variance between one delivery and the next depends on who took it on, and nobody can distinguish an improvement due to method from an improvement due to one unusually scrupulous person. The floor is not there to hold good people back. It is there to make their contribution visible, which otherwise blends into the noise.
What No Machine Can Declare Valid
There is one last step, and it picks up where I started a few days ago writing about the level you can’t delegate. This whole architecture of automated checks has a structural limit, and the limit is that none of them can declare its own output valid. A machine can inventory, can generate, can flag a gap, can block a completion claim. What it cannot do is mark its own work as accepted, compliant or validated, and not because that is technically impossible but because that gesture holds the only thing a legal system recognises, a person who answers for it. On this, European regulation stopped being an opinion some time ago and started drafting the org chart on our behalf.
The value is not in producing more artifacts, it is in producing evidence instead of confidence. A report that states, for every item, present, gap or not applicable, each with a pointer or a rationale, can be used in an adversarial proceeding. A report saying everything looks fine cannot, and the difference between the two is not tone, it is admissibility.
That is why I put the baseline I use in the open under an MIT licence: thirteen procedures plus one check that, at the end of a session, blocks the completion claim when a source file changed after the last passing test. There are tests verifying the procedures, 1,111 checks across nineteen suites, and a fixture repository built to fail, on purpose, every surface the harness audit inspects, because scaffolding whose soundness is asserted rather than shown falls squarely into the class of failures it is meant to prevent. It lives at oltrematica.github.io/oltrematica-skills, and anyone who wants to can dispute it line by line, which is the only serious way to offer something like this.
Two Questions for the Board
But the part a board needs is not the repository. It is two questions to put on the agenda of the next management meeting, and their usefulness lies in the fact that answering them requires no investigation.
The first: who, in this company, is required to verify that our scaffolding still works, and with what proof. The second: of what we currently call done, how much is held together by a constraint and how much by an intention. If the first has no name against it and the second has no check against it, this is not a tooling problem, and no purchase will solve it.
Key takeaways
There is a class of failures that produces no errors: an instruction file grown past the agent’s attention span, an automation description that has drifted until it fires on everything. There is no stack trace for scaffolding coming apart, and with no error to look at the team blames the model.
What degrades quietly is exactly what nobody is required to check. Application code has a defence against silent decay, and the defence is a test that gets angry. The instructions we steer agents with inherited the status of personal configuration: no owner, no changelog, no maintenance.
A definition of “done” written in prose will always compete with a delivery that has a date, and always lose. A deterministic check that blocks the completion claim when a source file changed after the last passing test is not stricter: it is a constraint instead of an intention, and under pressure only the constraint survives.
Training and baseline are two distinct investments. Training produces results above a floor, and in a team of ten the floor does not exist by default: it exists if somebody builds it, versions it and maintains it. Without that floor below, the variance between two deliveries depends on who happened to take them on.
No automated check can declare its own output valid, and not for technical reasons: that gesture holds the only thing a legal system recognises, a person who answers for it. The value is not in producing more artifacts, it is in producing evidence instead of confidence.
Questions & answers
Why is it healthy to distrust anyone proposing to revise the process?
Because most of the time what they are proposing is to slow down, and they propose it at the worst possible moment, when a team has finally found a rhythm. Every methodology of the last thirty years arrived as quality and was delivered as bureaucracy, and a good share of the AI governance pages written in the last two years exists so that somebody can tell a board the topic is covered. The suspicion is healthy. It simply does not cover the failures that produce no errors.
Why does an agent seem to get worse after a few months?
Almost never because the vendor degraded the model. More often the project’s instruction file has grown past the point where the agent can hold all of it, and it started quietly dropping the half that mattered; or the description of an automated procedure drifted through rewording after rewording until it fires on anything, which amounts to firing on nothing. Neither failure throws an exception, and with no error in sight the drop gets attributed to the tool instead of the scaffolding.
What difference does a deterministic check make compared with a written rule?
A rule in prose, say a definition of “done” in a page of internal conventions, is honoured when there is time and skipped when there isn’t, which is when it matters. Not out of negligence: because it will always compete, and always lose, against a delivery that has a date. A check that blocks the completion claim when a source file changed after the last passing test is not a difference of rigour, it is the difference between an intention and a constraint. Under pressure only the second survives, and the cost that looks saved moves to review, rework and the reopening of tasks already marked closed.
Why does nobody regenerate the dependency inventory mid-sprint?
Not out of ignorance. No company that has read a public tender in the last three years doubts that a software bill of materials matters. The point is that nothing forces the moment to exist: there is a release to ship, and regenerating the inventory competes with that release for somebody’s attention and loses, every time, until the day an auditor asks and it isn’t there. The difference from the other failures is that here the clock is public: Cyber Resilience Act vulnerability-reporting obligations apply from 11 September 2026, the regulation applies in full from 11 December 2027, and the European Accessibility Act has been in force since 28 June 2025.
Why should training and the baseline be funded separately?
Because they answer different questions. Training raises people, widens what they can see, moves someone from execution to specification and verification. But it produces results above a floor, and in a team of ten that floor does not exist by default: it exists only if somebody builds it, puts it under version control and treats it as a maintained artifact rather than a folder of notes. Without that floor below, everyone works at their own level of rigour and nobody can tell an improvement due to method from an improvement due to one unusually scrupulous person.