Commentary

Your AI Pilot Is Successful Because Nobody Has to Live With It

A successful demonstration can leave the hardest deployment questions unanswered. Leaders should demand a pilot that tests who will operate the system, absorb its costs and own its consequences.

Imagine a customer service pilot that everyone wants to call a success.

An AI assistant drafts replies. A small group of experienced employees reviews them. The project team resolves difficult cases. The sponsor sees faster drafting and approves the next phase.

Now imagine the same system running on an ordinary Monday. The experienced reviewer is on vacation. A customer disputes an answer. The system needs information that another department controls. The project team has moved on. Someone has to decide whether the answer can be sent, whether the customer deserves a correction and whether the service should keep running.

This is a hypothetical example. But it exposes the question that should govern the investment: did the pilot test the proposed operating model, or did it test what the technology could do with unusually generous support?

My position is straightforward. A pilot intended to justify deployment should earn approval under conditions that resemble the work people will actually inherit. If its success depends on arrangements nobody has agreed to sustain, its business case is unfinished.

The pilot has an invisible support team

Temporary support can make a pilot useful. Experts can diagnose problems, help employees learn and establish whether an application has potential. Early experimentation deserves room to work.

The trouble begins when temporary support disappears from the account of success.

Consider what the sponsor is being asked to approve. A reply appears in seconds, but an experienced employee checks its accuracy. Another person retrieves a missing record. A project lead handles the exception. The demonstration credits the system for the completed answer while leaving the surrounding effort outside the claim.

That effort is part of the service. Someone must fund it, staff it and make it reliable. If deployment requires a dedicated review team, that may still be a good investment. The decision depends on the whole arrangement, including the work the assistant creates.

The same applies to authority. A reviewer can reject an answer, but can that person pause the system? Does an operations manager have access to investigate a dispute? Can a frontline employee return to the previous process when the assistant becomes unavailable? Naming a human somewhere in the workflow does not settle those questions.

An executive should be wary of the phrase “we will address that when we scale.” It can mean that the most consequential part of the proposal has been deferred until after the organization feels committed to it.

Approval means someone accepts the consequences

A technical result answers a bounded question about performance. Deployment approval commits people, money and authority to an ongoing service.

Those commitments deserve to be tested together.

For the customer service example, the business needs to establish who may send a reply, which replies require additional review, who handles a disputed answer and how the customer obtains a remedy. The reviewer needs time, relevant records and the authority to intervene. The manager needs a budget that includes those responsibilities.

These are operating decisions with consequences for speed and cost. A requirement to review every reply may preserve control while consuming much of the expected benefit. A narrower review policy may improve throughput while creating a different exposure. The pilot should produce evidence that helps leaders choose between those arrangements.

NIST's AI Risk Management Framework gives this responsibility a formal place: its GOVERN function addresses executive responsibility, documented roles and human oversight. That supports the need to assign responsibility; it does not prove that a particular pilot will succeed or establish how much oversight an organization needs. NIST AI RMF Core.

My argument goes a step further. The people who will own the service should have a meaningful say in what counts as a successful pilot. A sponsor can support the investment. The operating team must be able to explain how it will live with the result.

Four commitments a deployment pilot should test: authority to act, capacity to review, ownership of exceptions and resources to sustain the service. Original Eldris conceptual guidance.

Original Eldris conceptual guidance. These commitments are questions to test, not a scoring system or a guarantee of deployment success.

Exploration deserves freedom. Deployment needs an owner.

The strongest objection is that this approach could burden experimentation with production requirements before anyone knows whether the technology is useful.

That objection is right about exploratory work. A team learning whether an assistant can handle a task should be able to run a small, bounded experiment. It need not build an entire operating model before discovering whether there is anything worth operating.

But the claim attached to that experiment matters. “This application merits further testing” is a defensible outcome. “This business unit should deploy it” requires additional evidence.

Leaders should make that distinction explicit when funding the work. An exploratory experiment buys learning. A deployment pilot tests a proposed service. Confusing the two allows evidence about technical possibility to acquire authority over a business decision it has not yet answered.

The solution is proportionality. A low-consequence drafting aid needs a lighter test than a system authorized to make commitments to customers. The ambition of the deployment should determine the ambition of the evidence.

Make ordinary conditions part of the test

Before the next pilot, ask the future operating owner to specify the conditions under which the service would be acceptable. Then design the evaluation around them.

Keep an account of assistance from experts and the project team. Include their review, retrieval and exception work in the economics. If the service depends on specialist intervention, establish whether that intervention can be supplied at the expected workload.

Test the routes for disagreement and interruption. An employee should be able to challenge an output and know what happens next. The team should know how to continue essential work if the service is paused. Exercises can test those arrangements safely; nobody needs to expose customers to an avoidable failure merely to prove that a response procedure exists.

For a multinational deployment, repeat the assessment where operating conditions materially differ. A headquarters team with direct access to experts may offer support that another business unit cannot reproduce. Language, records, approval practices and local responsibilities belong in the deployment decision.

Finally, allow the pilot to recommend a narrower service. A finding that the assistant can draft replies but should not send them is useful. So is a finding that one category of cases is worth automating while another should stay with an experienced employee. Scope is a management choice, and a pilot should help make it.

The outcome worth celebrating is a service whose future owner can defend its scope, resource needs and consequences. A polished demonstration may start that conversation. It cannot finish it.

Before approving deployment, ask the people who will inherit the system to show how they will operate it on an ordinary Monday. Their answer belongs in the definition of success.

More insights

Have a harder version of this question?

The writing is the general case. An advisory engagement is the specific one.