Practice
The Pilot That Worked Too Well
Pilots succeed and deployments fail. It is not the tool that changes, it is the environment that carried it.
Let us follow a pilot, from its start to its deceptive end. A legal department chooses a promising tool and entrusts it to a small team for three months. Five people, motivated, who talk to each other every day. They test the tool on a restricted perimeter, adjust their usage as they go, pass on best practices orally, correct each other. At the end of the quarter, the assessment is excellent: measurable time savings, real enthusiasm, obvious return on investment. The decision follows, logical: generalize.
Six months after deployment at scale, the picture is entirely different. Usage did not explode as the pilot led one to hope; it stagnated, then receded. Users use the tool for a few simple tasks and return to their traditional methods for the rest. The tool is still there, under license, but it is no longer central. No one quite understands why, since it is exactly the same tool that had worked so well in pilot.
And that is precisely where the answer hides. The tool is the same; the environment has completely changed. In pilot, coherence was not held by the tool, it was held by the five people around it. Context was shared naturally because everyone worked on the same perimeter. Coordination was direct because it sufficed to turn to one’s neighbor. Usage rules were informal because a small team can adjust without procedure. The tool bathed in an environment of coherence the humans produced around it, without even realizing it.
In pilot, it is not the tools that hold coherence. It is the humans around them.
The human net that vanishes at scale
When you deploy this same tool at the scale of an entire firm or department, this human net vanishes. Matters multiply and diverge. Users no longer know each other, no longer coordinate directly. The informal rules that sufficed for five become unmanageable for two hundred. Context ceases to be shared spontaneously. The tool has not changed by a byte, but the environment that made it useful has ceased to exist. One generalized the tool without being able to generalize what made it work.
One must see that this human net was invisible precisely because it worked. No one, in pilot, said to themselves “I am holding the team’s coherence”; each did it naturally, without naming it, like breathing. It is this invisibility that traps the organization: what it saw succeed in pilot was not the tool alone, it was the tool wrapped in a human work of coherence so natural it was never perceived as an ingredient of success. One generalizes what one saw, the tool, leaving behind what one did not see, the wrapping.
The pilot does not test the tool. It tests the tool surrounded by humans who compensate for its gaps.
What small size offers for free
To understand why the pilot systematically deceives, one must see all that a team’s small size provides without billing. Five people who talk every day share a context effortlessly: what one knows, the others learn in passing, in a hallway conversation, a remark said aloud. They coordinate without procedure: it suffices to turn to one’s neighbor. They correct each other without a device: a drift is seen and flagged immediately. All this is free, instant, and perfectly invisible, because produced by the mere proximity of the people.
This free provision is a methodological trap, because it distorts the measurement. When one evaluates a pilot, one attributes to the tool all the observed gain, when a considerable part of that gain comes from the free environment the small size provided. One measures the tool plus its wrapping, and attributes the result to the tool alone. It is an imputation error: one credits the visible component with a result produced by the invisible component, and one makes a generalization decision on the strength of this false imputation.
The pilot therefore suffers from a structural bias no rigor of execution corrects. One can run the pilot with the greatest care, measure the gains precisely, document the usages: as long as it takes place at small scale, it will keep measuring the tool in an environment that will no longer exist at scale. It is not a flaw of method that a better method would fix; it is a limit inherent to the pilot device itself, which by construction tests in non-reproducible conditions.
The four functions the pilot hides
One can push the analysis further and name exactly what the pilot hides. In a small team, four invisible functions are permanently ensured by the people. Someone always knows where the matter stands. Someone remembers what was decided last week. Someone spots that a usage drifts from the common rule. Someone keeps in mind who is entitled to see what. These four functions are written nowhere, they cost nothing apparent, and they are nonetheless the base that makes the tool useful. The pilot does not test them, because they are provided free by the small size of the team.
At scale, these four functions do not disappear: they become impossible to ensure by hand. No one can any longer know where each of three hundred matters stands, nor remember all the teams’ decisions, nor monitor all usages, nor hold the permissions map in their head. What was free in pilot becomes, at scale, either prohibitively expensive or simply absent. And it is their absence, not a defect of the tool, that makes usage drop. The tool finds itself naked, stripped of the wrapping that made it useful, and its raw performance, alone, does not suffice to hold real work.
This is why the deployment’s failure is not a failure of the tool, and changing tools will change nothing. What the deployment lacks is an organizational layer that takes charge, at scale, of what the small team ensured by hand: shared context, methodological coherence, usage coordination, traceability, governance. At five, these functions can rest on the people. At two hundred, they must be carried by the architecture, or they are carried by no one.
At five, coherence rests on the people. At two hundred, it must be carried by the architecture, or by no one.
There is a reason these four functions do not simply let themselves be recreated by procedure, once at scale. One might think it suffices to write rules, name owners, formalize what was informal. But what the five people did was continuous, contextual, adaptive: they did not follow a rule, they judged in situation. Turning that into procedure produces a heavy bureaucracy that captures poorly what it claims to replace, and that no one follows at two hundred. These functions do not re-formalize; they re-tool, by entrusting to a layer what procedure will never be able to hold.
Why the loop repeats for years
There is a psychological detail that worsens the misunderstanding, and it is worth naming. The successful pilot creates an expectation, almost an implicit promise made to the organization: if five people saved so much time, three hundred will save proportionally more. This extrapolation seems obvious, and it is false, because it supposes the gain comes from the tool alone, when it came from the tool plus the environment of coherence the five people produced for free. One extrapolates the numerator while forgetting the denominator.
This is why the deployment’s failure is so often experienced as a betrayal, and so rarely understood. Leaders saw the pilot with their own eyes; they know the tool works, since they observed it. When usage collapses at scale, they look for a culprit: a training defect, a resistance of the teams, a bad tool choice. They change tools, relaunch a pilot, which succeeds again, redeploy, and fail again. The loop can last years, because the real cause, the absence of a layer, is never named.
Escaping this loop supposes a change of view on what a pilot is. A pilot does not prove a tool will hold at scale; it proves only that it is good in an environment where humans hold coherence in its place. It is useful but partial information, and confusing it with a proof of deployment is the most costly error of the current market. The good pilot is not the one that tests the tool alone, but the one that tests the tool under the orchestration layer that will accompany it at scale.
A pilot does not prove a tool will hold at scale. It proves it is good when humans hold coherence in its place.
Testing the right object
This is exactly the role for which MAX was designed: a Legal Semantic Layer that maintains at large scale what pilots hold at small scale informally. A layer that transforms the passage from pilot to deployment, today experienced as an unexplained regression, into a mastered change of regime. The pilot succeeded because humans held coherence; the deployment succeeds when a layer holds it in their place. It is not a supplement to the tool, it is what was missing between the pilot and scale.
This shift changes the very way of conducting a pilot. A useful pilot is no longer the one that puts five motivated people around a tool to see if they use it well, since they will always use it well at five. It is the one that tests the tool in the conditions of real deployment: without the human net, under the layer that will have to hold coherence at scale. Such a pilot predicts something; the other predicts only its own success, in conditions that will never recur.
The next generation of organizations will therefore no longer judge a tool on its pilot performance, which proves almost nothing, but on its ability to hold once the orchestration layer is in place. The pilot will remain useful for evaluating a brick; it will cease to be confused with a proof of deployment. And the day this distinction is acquired, the loop of successful pilots and failed deployments, which costs the market years and entire budgets, will finally stop repeating.
The tools did not fail at scale. The architecture that should have surrounded them never existed.