Research

THE AI PILOT WORKED. TWO YEARS LATER IT IS STILL A PILOT

Mike Gault·
THE AI PILOT WORKED. TWO YEARS LATER IT IS STILL A PILOT
The pilot problem with AI agents

Almost every health system has an AI success story that never left the department it started in. An ambient documentation trial in one clinic. A triage assistant on one ward. A scheduling agent handling one team's inbox. The results were good enough to present at board level, and then nothing happened.

The trouble is - small deployments work because someone is watching, in a controlled environment. The problems hit when you connect with the messy real world

A pilot succeeds because a named person is paying attention to it. There is a clinical lead reviewing output, a project manager who knows every case the system has touched, and a small enough patient group that anything strange surfaces within a day.

That arrangement is exactly what scaling removes. Running the same tool across forty sites means nobody is watching in the way they were during the trial. The oversight has to be described, written down, and built into the system itself. This is where most programmes stop, and it has nothing to do with model quality.

Killer questions for AI production readiness

Three questions come up in every scale-up review, and they are usually asked by legal, by the CISO, and by the chief medical officer in that order.

  • Who is accountable when the agent acts and no human reviews the decision before it happens?
  • What does it look like when the system is wrong at forty sites at once rather than one?
  • How do you show, six months later, what the system actually did and why?

None of these are answered by a better model. They are answered by knowing where the system runs, what it is allowed to do, and what record it leaves behind.

Most AI programmes in healthcare have a clear vision and no description of the machinery underneath it, which is why they keep passing the pilot stage and failing the production one.

This is not a healthcare-specific failure

Gartner's Max Goss put a figure on how widespread the gap is in April 2026: the average Fortune 500 company is expected to run more than 150,000 AI agents by 2028, up from fewer than fifteen today, and only 13% have adequate governance in place.

Healthcare is not behind the rest of the economy here. It is simply the sector where the consequences of getting it wrong arrive fastest and are hardest to reverse.

The regulatory timetable

The implementation schedule for the EU AI Act has shifted, and it is tempting to read that as breathing room. It is not. The obligations fall on the organisation deploying the system, not on the vendor supplying it, and that allocation of responsibility does not depend on a commencement date. A hospital that cannot describe what its agents are permitted to do has the same exposure it had before the timetable changed. The same is true under existing data protection and clinical governance rules, which never went anywhere.

What has changed is that the deadline is no longer doing the work of forcing an answer. Organisations that were waiting for a date now have to decide on their own terms whether they want to be running agents at scale in two years, and if so, what has to exist before that is possible.

Making AI pilots production ready in healthcare

The useful shift is to stop treating the pilot as a test of whether the model is good enough, and start treating it as a test of whether the deployment can be described.

Three things need to be true before a trial can become an estate-wide rollout.

  • Every agent needs an identity that can be traced to a person or a team and withdrawn without touching anything else.
  • Every agent needs limits that hold whether or not the model behaves as expected, enforced by the surrounding system rather than by instructions written into a prompt.
  • And every action the agent takes needs to produce a record as it happens, rather than one assembled afterwards from whatever the logs happened to capture.

A pilot that can demonstrate those three properties is a pilot that legal can sign off at scale. A pilot that cannot is a demonstration, however good the clinical results.

Architecture requirements

Existing approaches each solve part of the problem - system prompts control behaviour, containers isolate execution, policy engines enforce rules. But none deliver the full stack: multi-user isolation, swappable components, audit, and sandboxing in a single runtime.

Our product; Unicity AOS gives you a pluggable brain in a padded room with a mail slot - the loop is isolated from everything it controls. It can't bypass the sandbox, skip the audit, or access tools it hasn't been granted.

Next steps

The practical next step is to take the last successful pilot in your organisation and ask what would have to be added to run it unattended across every site. The gap that comes back is the actual project.

To get started why not see an AI agent hit a regulated boundary and get stopped

Watch an agent attempt a prescription, across-bordertransfer, an unconsented store - and get intercepted in-path, live, on your own infrastructure.

BOOK A LIVE DEMO