For a few years we told ourselves that every company was sitting on a gold mine called proprietary data. The formula was simple and looked unassailable: general-purpose model plus data nobody else owns equals better vertical AI. All you had to do was wait for models to get smart enough, plug them into your archives, and let twenty years of accumulated work do the rest.
It is a reassuring metaphor. I suspect it is also wrong, or at least incomplete in the one place that matters.
An archive contains what an institution has done. Not necessarily what it has learned.
A Thesis Worth Taking Seriously
Before dismantling it, the proprietary-data thesis deserves respect, because the argument is solid. Foundation models learn from enormous amounts of public or licensed information. As model quality converges, whatever a competitor cannot download or buy naturally becomes the source of differentiation. A hospital holds clinical data OpenAI does not hold. A bank sees transactions Anthropic never sees. A law firm keeps opinions and precedents that are not on the internet.
And this data is hard to replicate for a structural reason: it is the byproduct of years of activity. A new competitor can buy GPUs, hire researchers, license a foundation model. It cannot recreate in a quarter twenty years of interactions with real clients, real failures, decisions taken in the real world.
This summer’s most quoted case seems to confirm all of it. On August 24 Thomson Reuters announced Thomson, a proprietary model built on an open source base with its own mid-training and post-training techniques, specialised with the Westlaw, Practical Law, Checkpoint and Reuters corpus, at a declared investment of $40 million in talent and compute. A tiny fraction of what a frontier model costs. It looks like the perfect demonstration of the data moat: they owned content that OpenAI, Google and Anthropic do not have on the same terms, and they built on it.
Then you read the most interesting detail of the announcement, and it is the one that makes the whole interpretation insufficient.
What Thomson Reuters Did Not Do
Thomson Reuters states that, to date, it has used less than ten percent of its content in the model’s continued pre-training. And it explicitly attributes the results not to access to documents, but to the combination of authoritative content, domain specialists, professional tools, training objectives and evaluations built with hundreds of subject matter experts. The experts did not supply examples of correct answers: they judged outputs and wrote rubrics enumerating what a correct answer must necessarily contain; the failure modes came from thousands of internal experts using the model on their hardest problems, errors reported by people qualified to diagnose them.
Stay on this point, because it is the heart of the essay. The company with one of the most valuable professional corpora on the planet did not pour the corpus into the model. It spent the hard part of the project deciding what learning correctly from that corpus should mean. Ninety percent of the content stayed out, and nobody seems to consider that a waste.
The “model plus data” formula confuses the raw material with the industrial process that gives it value. Two organisations can hold comparable amounts of data and get radically different results from it. One can have millions of documents without knowing which are reliable, thousands of tickets without knowing which represent meaningful exceptions, decades of decisions without having kept the reasons they were taken.
The deepest competitive moat may not be proprietary data. It may be proprietary judgment. And judgment, unlike data, does not live in a database. It lives in the culture of the institution.
We Archived the Outputs, Not the Decisions
Picture two law firms. Same sector, twenty years of contracts, roughly the same number of tokens. The first keeps the signed versions. The second also keeps the drafts, the rejected clauses, the partners’ comments, the alternatives considered, the reasons an apparently standard wording was not used, the incidents that happened after signature.
Statistically both own “proprietary data”. But only the second owns the trace of judgment. The signed contract records what we decided; the history of the negotiation records why we decided this rather than the available alternatives. And it is almost always the second element that carries the professional advantage.
Modern organisations are extraordinarily good at keeping artifacts: documents, tickets, commits, contracts, minutes. They are much worse at keeping the space of alternatives that was crossed to produce them. We see the final merge, not the three architectures a senior mentally discarded before proposing it. We see the clause, not the risk that killed the more convenient one. We see the accepted quote, not the informal assessment through which someone understood that an apparently small requirement would blow up the project.
The most important asset of the organisation may be precisely what never became data.
The Material Inside “It Depends”
Michael Polanyi gave this problem its most durable name: we know more than we can tell. A significant part of professional competence is tacit. A senior looks at a technical proposal and feels something is off. A lawyer recognises that a theoretically correct clause is dangerous in that context. A doctor reads a combination of signals that, in isolation, would say nothing.
Ask “which rule did you apply?” and the answer is often “it depends”. In the rhetoric of automation, “it depends” is the enemy: the sign that a process cannot be formalised. I believe the opposite. Behind that “it depends” sits a structure of exceptions, priorities and trade-offs the expert has internalised over years of cases. It is the most valuable material the organisation owns, and it is exactly what no SharePoint export contains.
Truly differentiated vertical AI may be the technology through which an organisation gradually makes its own “it depends” explicit. Not by interviewing the experts, which is almost always the worst method: tacit knowledge does not verbalise in the abstract. By making them judge. Two answers, two architectures, two clauses: which one do you pick? Why? And if I change this detail of the case? The point where the expert changes their mind reveals the boundary of the rule. The model can generate the alternatives; the expert judges them; the system captures not the preference but the reason, and the conditions under which the preference would flip.
Frequency Is Not Authority
There is an error symmetrical to ignoring your own archive: trusting it too much. Suppose you have a million support tickets. It looks like a magnificent dataset. It contains tickets closed well, terrible workarounds that became habits, operator mistakes, temporary fixes that outlived the constraints justifying them, behaviours that worked once by accident.
Training indiscriminately on that corpus means converting historical frequency into epistemic authority. Institutional culture exists precisely to prevent this: to say “this happened often, but it was not right”, or “this case occurred three times, and those three times describe the most important exception in our domain”. Data counts occurrences. Judgment assigns them meaning.
Machine learning has a precise name for the mechanism that defines what it means to be wrong: the loss function. Two models can observe the same data and learn different behaviours if they are optimised against different objectives. Organisations work the same way. For a software house, a cosmetic bug, a vulnerability, an accessibility regression and a delivery delay do not carry the same cost, and two companies price the same events differently. Institutional culture is, in large part, the organisation’s accumulated loss function: it says which mistakes are unforgivable, which trade-offs acceptable, which risk deserves escalation, where to be conservative and where to experiment.
Data can be bought, and it is becoming less exclusive: datasets get licensed, synthetic data fills gaps, retrieval uses content without baking it into weights. A good organisational loss function takes years, because it is made of training, selection, conflict, correction and memory. A competitor can buy ten million court rulings; it does not buy the ability to tell which argument is legally correct but strategically suicidal.
Whoever Owns the Benchmark Owns the Definition of “Good”
Hence the most technical consequence of the argument. In these years we have given enormous weight to training, fine-tuning, RAG, vector databases. If models keep improving and becoming interchangeable, the most durable part of the stack may be something else: the evaluation system.
Not “our model can answer tax questions”, but “we own two thousand cases that represent exactly what giving a correct tax answer means in the situations that matter to our clients”. Each case with its context, the expected answer, the alternatives that look correct and are not, the severity of each error, the acceptable sources, the conditions under which you stop and call a human.
At that point you can swap GPT for Claude, Claude for an open-weight model, that model for the next one. The advantage remains, because you know how to measure which engine deserves to enter your process. Frontier models are already very good at producing plausible results; the vertical advantage emerges in the regions where plausible and acceptable diverge, and that is where the expert is needed.
Thomson Reuters is moving exactly in this direction with CoCoBench: more than a thousand attorney-authored tasks reflecting what practitioners actually do, scored against reference answers drafted and reviewed by attorneys, over fifteen thousand hours of work by more than a hundred legal experts. One scheduling detail says it all: CoCoBench is from May, the model from August. The measure arrived before the engine. And the observation that comes with the benchmark applies far beyond legal: when you change the level at which you evaluate, from single task to workflow, what counts as good changes. Benchmarks look like neutral instruments and never are: each one embeds an implicit theory of what counts. Measure a coding agent on the amount of code produced and you reward one thing; measure it on its ability to recognise when it lacks sufficient information and you reward another. Building proprietary evals is an organisational act disguised as a technical one: it formalises what the institution considers good judgment.
And the CoCounsel platform itself remains deliberately multi-model: it applies Thomson where it delivers the clearest advantage and other models where they are better. As for the base, they write that they have already changed the root model several times and intend to keep doing so as the open-weight frontier moves. The implicit message is not “we managed to build our own GPT”. It is: we can change the engine without losing our definition of correctness.
Culture Can Be Wrong
Here comes the serious objection, and it deserves space because it is well founded. Celebrating institutional culture as a competitive advantage risks turning into a conservative defence of the status quo. Organisations also accumulate prejudices, obsolete procedures, pathological hierarchies, apparent consensus, rules born from incidents that no longer matter, mistakes carefully handed down from senior to junior. Encode all of that into evals and training data and you are not preserving wisdom: you are industrialising tradition.
The advantage does not consist in teaching the AI “how we do things here”. It consists in knowing which parts of how we do things deserve to be preserved and which must remain contestable.
The most concrete defence I know is preserving dissent. If we build datasets only from final decisions, we erase a fundamental part of the organisation’s intelligence: disagreement. Three seniors review an architecture, two approve it, one considers it risky for a precise reason. Traditional knowledge management archives “architecture approved”. A system that manages judgment should archive “approved two to one; risk X considered and deemed acceptable because Y”. Six months later X happens, and the minority dissent becomes the most valuable information in the archive. A computable culture without a memory of dissent turns historical consensus into truth.
For the same reason, counterfactuals are worth keeping. “We chose PostgreSQL” is worth little. “We chose PostgreSQL because requirement X demanded Y; without X we would have preferred SQLite” is worth a lot, because it lets a future agent understand when it must not imitate the past. That is the difference between memory and intelligence, and it is why naive retrieval is dangerous: it finds the most similar precedent and applies it, while an intelligent organisation knows how to recognise when the new case is different enough to make the precedent useless.
And like everything that can be wrong, codified judgment must be versioned. Who may modify a rubric, who decides a failure mode is no longer relevant, who changes an escalation threshold, how often the evals themselves get challenged: these are governance decisions, not implementation details. A healthy judgment archive must be able to say “this was our conviction in 2026, these incidents put it in crisis, since 2028 we use a different rule”. Institutional memory exists to change your mind better, not to make changing it impossible.
The Apprenticeship Paradox
There is, however, a risk deeper than all the others, and it is a paradox. Professional culture does not arise spontaneously: it forms through apprenticeship. The junior reads, compares, makes mistakes, receives corrections, watches how a senior handles an ambiguous case, learns which details deserve attention. Much of this path coincides with the activities AI is making cheapest to automate.
Thomson Reuters’ Future of Professionals 2026 report puts numbers on this fear: 48% of professionals fear a negative impact of AI on the development of independent judgment, and legal respondents estimate the path to trusted professional judgment could lengthen by nearly two years (1.7, to be precise; tax professionals, it must be said, expect the opposite). The fear the report records is that the craft gets hollowed out: that judgment, expertise and the human relationship stop being valued. I would call it hollowing out the judgment pipeline. By eliminating the work through which professionals used to learn, we make the present organisation more efficient while consuming the mechanism that produces its future experts.
I wrote about this when I argued that the level of judgment cannot be delegated: here the paradox becomes strategic. The best vertical AI requires mature institutional culture; AI itself erodes the social process through which that culture is transmitted. A company rational about the quarter eliminates every activity the model performs better and cheaper. A company rational about the decade asks which apparently inefficient activities are in fact the laboratory where the next senior is formed.
It seems useful to name what gets consumed: epistemic capital. The accumulated ability to tell good decisions from bad ones, recognise exceptions, diagnose errors, transmit these criteria to the next generation. When we hand junior work to AI we gain immediate productivity and, perhaps, eat into this capital. The problem is that the bill never shows up in the accounts: the indicators improve, more output, fewer hours, better margins. The damage becomes visible years later, when you discover you have plenty of operators able to use the system and very few people able to notice the system is wrong.
The countermove is not slowing adoption. It is changing what seniors and juniors produce. The senior stops being only the person who solves the hard cases and becomes the person who defines the evals, identifies failure modes, selects edge cases, makes thresholds explicit: their work produces value twice, because it solves the present problem and improves the system that will solve the next ones. And the junior does not spend less time thinking: they spend less time mechanically producing the first draft and more time evaluating outputs, comparing alternatives, explaining errors, watching corrections. Explicit apprenticeship of judgment instead of production. The time of your experts spent making their judgment transferable is the real AI investment of many companies; calling it epistemic capex helps defend it in a boardroom, where it otherwise shows up as overhead.
There is an almost amusing corollary: the company that documents its own mistakes best might win. An honestly analysed incident produces an eval, a failure mode, a test, a precedent. Two companies make the same mistake; one fixes it and forgets, the other turns it into an artifact that makes it hard for any future agent to repeat it the same way. The second has converted a loss into institutional capital. Honest post-mortems are about to become one of the most valuable training materials of the agentic era, not least because they are exactly the kind of knowledge the public internet does not have.
The Judgment Stack
To keep the thesis from staying gaseous, the structure I would propose to an organisation has five layers.
The first is the evidence layer: the data and the sources. Documents, code, regulation, telemetry, precedents. It is the layer everyone talks about, and the least differentiating.
The second is the memory layer: what the organisation has done. Decision records, incidents, outcomes, the history of choices with their context.
The third is the judgment layer: how it tells good from bad. Rubrics, evals, failure modes, thresholds, exceptions, preserved dissent. This is where the advantage lives.
The fourth is the governance layer: who may change the third. Ownership, versioning, review, deprecation.
The fifth, deliberately last, is the execution layer: the models and the agents. The engine executes an already structured culture; it is not the place from which that culture should magically emerge.
Framed this way, even procurement changes. The question stops being “which LLM is best?” and becomes “which engine best executes our Judgment Stack?”. The organisation stops adapting the way it works to the quirks of this season’s model and puts models in competition against its own definition of value.
The Sovereignty of Criteria
This architecture has a consequence I care about, because it reverses a debate Europe conducts badly. If the advantage sits in the judgment layer, a company can use open-weight foundation models, open source frameworks, open standards and portable infrastructure without giving up its differentiation. On the contrary: the more the replaceable component is open and standardised, the more the specific asset concentrates where it is defensible, in the evals, the policies, the judgment data.
I wrote that the next lock-in will not hold your data but your accumulated state: this is the constructive face of the same argument. Sovereignty does not consist in owning the intelligence. It consists in owning the criteria by which we decide whether that intelligence deserves trust. An organisation that can say “we can change the model without losing what we know” is more sovereign than one that trained its own model and cannot evaluate it.
It also holds at scale. Europe is unlikely to win a race played on GPUs, capex and maximum-scale foundation models. But it holds sectors with enormous institutional density: healthcare, industry, engineering, public administration, pharmaceuticals, law. The interesting strategic question is not “how do we build a European OpenAI?”. It is: how do we turn decades of European professional culture into verifiable, portable, sovereign systems? That is not a loser’s consolation. It is a different strategy, and probably a more plausible one.
Succession, Not Substitution
The most concrete version of this thesis I see in small and medium Italian companies, which own very little “big data”. No billions of events, no data lakes. They own something else: people who have done an extremely specific job for twenty-five years, procedures never fully documented, exceptions known to three people, clients who taught them where an apparently correct spec falls apart. We usually call it a weakness: too much knowledge in people’s heads. And that knowledge is leaving, not because of AI but because of demographics.
Here AI can do something different from automation: interrogate those people, put cases and alternatives in front of them, surface the exceptions, turn part of their judgment into institutional patrimony before it walks out the door with the next round of retirements. Not substitution: succession. It strikes me as a more European reading of AI, and frankly a more useful one, than the reading that only counts jobs to defend or to cut.
One limit remains, and I state it against my own enthusiasm: a fully executable culture would be a dead culture. Living cultures change because experts contest the rules, exceptions generate new rules, regulation and markets move. The function of AI is not to freeze past judgment. It is to make visible the point where the present stops resembling the past closely enough. The best system is not the one that replicates the institution perfectly: it is the one that can say “here the precedent is no longer enough”.
We spent the first years of generative AI looking for what companies owned and models did not, and the most immediate answer was data. It was reasonable. But a company does not become competent by accumulating what it has done: it becomes competent by building, over time, the distinction between what worked and what must not be repeated, between the rule and its exception, between a plausible answer and an answer someone is willing to take responsibility for. That distinction lives in people, in corrections, in conflicts, in remembered errors, in the things a senior sees and a junior does not yet.
Proprietary data tells the model what the institution has seen. Institutional culture teaches it what, out of everything it has seen, deserves to become judgment. The foundation model can become a commodity: American, European, open-weight, swapped every six months. The organisation keeps owning the thing that actually matters. Not the answers, but its own, permanently contestable definition of what a good answer means.
Key takeaways
Thomson Reuters used less than ten percent of its content in Thomson’s continued pre-training and credits the results to hundreds of subject matter experts: rubrics enumerating what a correct answer must contain, failure modes reported by people qualified to diagnose them. CoCoBench, the benchmark of more than a thousand attorney-authored tasks, dates from May: the measure arrived three months before the engine. The corpus was the raw material; the hard part was deciding what learning correctly from it should mean.
Institutional culture is the organisation’s accumulated loss function: it says which mistakes are unforgivable, which trade-offs acceptable, where historical frequency carries no authority. Archives keep outputs, not decisions: making judgment computable requires the trace of discarded alternatives, preserved dissent (approved two to one, risk X accepted because Y), counterfactuals that say when not to imitate the past, and versioning with its own governance, because culture can also be wrong and must remain contestable.
The strategic paradox: the best vertical AI requires mature judgment, but AI consumes the apprenticeship that produces it. 48% of professionals fear for the development of independent judgment and legal respondents estimate nearly two extra years to form it. The answer is to treat expert time spent on evals and rubrics as epistemic capex, redesign junior work as an apprenticeship of judgment, and put the model at the last layer of a Judgment Stack: sovereignty is being able to change the engine without losing your definition of correctness.
Sources
- Thomson Reuters Leverages its World-Class Data Assets to Launch Its Own Frontier Model, Thomson Reuters, 24 August 2026
- How we built Thomson, Thomson Reuters, 24 August 2026
- Thomson: a purpose-built foundation model for professionals, Thomson Reuters, 24 August 2026
- Why Legal AI Needs a New Standard: Inside Thomson Reuters CoCoBench, Thomson Reuters, 4 May 2026
- Future of Professionals Report 2026, Thomson Reuters, 2026 May 2026-06
- The Tacit Dimension, University of Chicago Press, 1 January 1970