Skip to content
The Accountable Firm

Chapter 14 — Start Small Without Isolating the Business

Skip to this unit’s text

Chapter 14 — Start Small Without Isolating the Business

The usual choice is a false binary. Either the company rolls AI out everywhere at once, or it keeps experimentation safely away from the work that matters. The first approach spreads untested changes through live operations. The second produces a polished demonstration that never changes the business.

There is a better starting point: run one bounded experiment beside an existing process while keeping the current delivery system intact. The purpose is not to create an innovation island. It is to learn, under real operating conditions, whether one result can be delivered differently and what the business would need to adopt that method.

Choose a Process, Not a Department or a Role

A role is too narrow for the first test. It can show that one person became faster, but not whether inputs, handoffs, review, and the final result changed with that person. A department is usually too broad. It contains several flows, customers, and measures, making it impossible to see which change produced which result.

Choose one process instead. It should have a recognizable trigger and endpoint, recurring work, a result another person can accept, and a consequence the business would notice if the process failed. Customer response, maintenance diagnosis, contract review, recruiting screening, and internal proposal preparation can qualify when the company can state the actual result and boundary.

Before starting, name five things: the process owner; the result; the AI-supported action; the human judgment that remains; and the route for exceptions. A process without these elements is not ready for a pilot. More technology will only make its ambiguity faster.

Run a Parallel Test That Respects the Core Business

The established process continues to serve customers, revenue, quality, and daily commitments. A small group tests an AI-enabled version of the same class of work within an agreed boundary. That parallel design creates a comparison a broad rollout cannot: a familiar delivery method beside a proposed one, without forcing the core business to absorb every early error.

The comparison is not a staffing exercise. It asks whether the delivery method can be redesigned while the result remains acceptable. If a pilot is treated as proof that people should be removed, people have every reason to protect the exception logic and practical context that the pilot needs. The test then loses access to the evidence that would make it trustworthy.

State the boundary plainly. The pilot is testing a process and its accountability design. Its interim evidence may support learning, workflow repair, and later capacity decisions; it is not a shortcut to a headcount conclusion. That is not a promise that roles will never change. It is a condition for people to bring real work, real failures, and real judgment into the test.

The boundary should identify the class of cases the pilot may handle, the cases that remain in the established process, the information the pilot may use, the decisions it may prepare but not make, and the conditions that require a person to take the work back. A customer-facing pilot may organize history and draft a response while reserving commitments, unusual remedies, and relationship-sensitive exceptions for a person. The purpose is not to make the pilot harmless. It is to make its risk understandable enough that the organization can learn from it.

The first operating boundary should also be visible to people outside the pilot. A sales colleague should know whether the new route may make a customer commitment. A manager should know whether a pilot result can enter a performance discussion. A receiving team should know whether an exception returns to the established process or is held for the next review. Vague boundaries do not preserve flexibility. They shift risk to the person nearest the customer or the output.

This is why the early pilot should use a parallel route rather than a hidden workaround. A hidden workaround can make an early dashboard look clean, but it gives the organization no way to compare the real effort, exception rate, or recovery work. Parallel work is slower at first because it makes the old and new methods visible together. That cost is not waste. It is the price of learning whether a faster execution layer has improved the complete result or merely moved the burden elsewhere.

The pilot should also have a defined exit from parallel work. At the beginning, write what evidence would justify moving a class of cases into ordinary operations and what evidence would send those cases back to the established route. This keeps the group from drifting into a permanent hybrid in which nobody knows which method is authoritative. A parallel test is a temporary design condition, not a new organizational layer.

Before work begins, the accountable executive should set a short operating charter. It needs five decisions: the result the pilot is trying to improve; the boundary of cases it may take; the human judgments that cannot be delegated; the evidence required at review; and the authority to pause or stop the experiment. The process owner runs the test inside those decisions. The review owner checks whether the work still meets the agreed condition.

Build the Learning and Adoption Interface

A pilot fails to travel when its only interface is the pilot team’s connection to a tool. It needs a second interface with the business that will eventually run the work. Build that interface before the pilot has declared success.

The people who may inherit the method should be able to see the proposed result, the work that has moved, the work that has become more demanding, and the judgments that remain human. They need an honest account of what the workflow will ask of them: what information must be supplied at the beginning; which exception now arrives earlier; what decision can no longer be made by habit; and what record the next person needs when the output is challenged.

Adoption is an operating issue, not a communications campaign. If the receiving team cannot explain why a rule exists, cannot find the context behind it, or has no channel to correct it, the method is not usable. It remains a demonstration maintained by its original enthusiasts. A sound handoff gives the receiving team a way to challenge a rule, add an exception, and return a correction to organizational memory without waiting for the pilot group to reappear.

The adoption interface also needs a place for disagreement. A receiving team may see a failure mode that the pilot group did not encounter, or discover that an apparently simple rule removes context that matters in daily work. That is not resistance to the pilot. It is operating evidence. Give the team a simple route to return a case, request a clarification, or propose a changed rule. If the only response is “the pilot already decided,” the method will be followed ceremonially and bypassed privately.

A transfer should include more than a process diagram. It should include the purpose of the changed step, the condition under which it should not be used, the person who can interpret a disputed result, the source of the relevant context, and the date on which the team can revisit the design. The first reusable handoff is not a technical document. It is a compact operating agreement.

The pilot group may coach the first operating cycle and help make tacit judgment explicit. It should not become the permanent interpreter of every case. If it does, the company has moved work to a specialist enclave instead of changing how the business operates. The receiving team should take part in accepting the handoff. Measure transfer by whether the receiving process owner can run an ordinary case, the review owner can identify a challengeable output, and the knowledge maintainer can locate and update the relevant rule.

Make the First Pilot an Operating Experiment

Run the pilot on a real business cycle, not only on prepared examples. AI may retrieve relevant context, organize incoming information, prepare a first draft, flag a routine pattern, or assemble material for human judgment. Keep the change narrow enough that the group can see what it altered and where it created a new burden.

The first question is not whether the model produced an impressive answer. It is whether the workflow made an operating problem visible. Which review standard cannot yet be stated? Which exception has no escalation route? Which judgment remains trapped in one colleague’s experience? Where has work become slower because the handoff moved rather than disappeared?

Those are useful results. A pilot that exposes an unclear boundary has done more for the organization than a demo that conceals one. Treat each material exception as design evidence. It may show that the workflow lacks context, the acceptance condition is vague, a named person has responsibility without authority, or the process boundary is wrong. The response may be to narrow the AI-supported action, add context, move the decision to the person who can make it, or stop delegating that class of work for now.

The same discipline applies to apparent success. A workflow that produces a fast first draft may still consume more review time, create a new coordination burden, or lead customers to receive inconsistent answers. Record where the work goes after AI has touched it. Does a manager now spend more time checking? Does a customer wait for a person to correct a commitment? Does a new exception go to an inbox nobody owns? A local productivity gain is not yet a changed operating result.

Do not compensate for an unclear design by adding review after review. A stack of confirmations can make the pilot look controlled while leaving nobody able to explain why the controls exist or which one can change the outcome. A good pilot reduces ambiguity. It does not cover ambiguity with paperwork.

Accept Each Stage for What It Proves

A first pilot has three jobs. The first is understanding: the group can explain what AI can and cannot do in the process, where human review remains necessary, and where the next bottleneck appears. The second is reproducibility: the group turns working knowledge into a process description, review rules, exception routes, decision records, and reusable examples. The third is embedding: the business can run the process with its own accountable roles and operating rhythm.

These are different tests. Understanding asks whether the organization has learned how the work changes. Reproducibility asks whether that learning can travel. Embedding asks whether the business can use it without a permanent escort. Do not collapse them into one adoption number.

The distinction matters because a pilot can pass one stage and fail another. A group may understand the workflow well but have no usable way to record exceptions. It may create a detailed method that no receiving team can run under normal time pressure. Or the method may travel, only to reveal that the business has not assigned someone authority to resolve a recurring conflict. Each failure has a different remedy. Treating all of them as “adoption problems” hides the actual design work.

Review the Experiment as a Business Decision

The review cadence should match the process, but its logic should remain stable. A short weekly discussion can surface an exception, a gap in context, or new rework. A more formal decision point should examine whether those observations change the method.

Bring the same questions to each review. What result did the process produce? Which work became easier, and which became harder? Where did a person exercise judgment that the workflow could not carry? What was corrected, and can the next person use that correction? Did an exception reveal a gap in authority, context, or acceptance?

Every review should end with a choice: continue within the boundary; revise a rule, input, or review point; prepare a bounded replication; or stop. Name the next accountable action and the evidence that will be considered at the next review. A pilot that continues because no one has decided otherwise becomes a shadow operation.

The review record should preserve dissent. If a reviewer believes the result is not ready, if a process owner believes the context is incomplete, or if a receiving team believes the proposed handoff shifts too much risk, record that judgment and the response. A meeting that records only agreement cannot teach the next team how the organization handled a genuine trade-off.

Keep the review materials proportionate. A one-page operating view is usually enough: intended result; current evidence; exceptions; changes to context, authority, or workflow; and the next decision. Long presentations are often a warning sign that the group has no shared result boundary. A useful review lets a decision-maker see not only whether a tool worked, but whether the organization has become able to own the consequence of using it.

Know When the Pilot Can Enter Daily Operations

Four conditions should be visible before an AI-enabled method moves from a bounded pilot into ordinary operations. The result is stable across ordinary work, not only an unusually successful run. The accountability chain is clear: process owner, review owner, acceptance condition, and escalation route are known. Another team could use the method without rebuilding it from scratch. And the business still has enough trust to provide the context and judgment the workflow needs.

Stop a pilot when the result cannot be made acceptable within the agreed boundary; when necessary judgment cannot be given adequate authority; when context cannot be used responsibly; or when the exceptions reveal a consequence the current organization cannot carry. Close a pilot when the method is stable enough to enter ordinary operations and the receiving business can carry the next cycle.

In either outcome, retain what was tried, what failed, what rule changed, and what another process should learn. A stopped pilot is not wasted work if it makes a hidden organizational limit visible early. The pilot is a bridge into the business only when it produces a method, a responsibility boundary, and learning the business can carry forward.

Before closing the first cycle, ask one final question: what would the next team need in order to avoid repeating the same first-month confusion? The answer may be a clearer result definition, a better source of context, a narrower AI action, a visible acceptance condition, or a stronger escalation route. Capture it. The organization does not become more capable because it has completed a pilot. It becomes more capable when the next pilot begins with the first pilot’s judgment already available.

For the same reason, avoid choosing a pilot that only produces internal artifacts. A proposal assistant, an intake process, or a knowledge workflow may be appropriate, but the group still needs to name the downstream result that gives the work meaning: a better decision, a more reliable response, a faster qualified handoff, or less rework in a consequential process. Without that result, the pilot can optimize a visible task while leaving the business unchanged.

At that point, the pilot has earned the next question: can the organization make the method repeatable without relying on its original team?