Skip to content
5 min read

A demo is not a process: what an AI pilot must prove

A demo does not prove the process improves. Here is how to define what an AI pilot must prove before you build.

Mauricio Zaffari

You watched an AI demo. It read a PDF, extracted the fields and filled the screen with a result in seconds. The decision in front of you now is concrete: commit budget to build, or commission something that proves the process actually improves first.

Do not commit build budget on demo evidence alone. Set the pilot criteria first, and use a document-extraction example to test when to continue, adjust or stop.

What the demo is worth as evidence

A demo runs on a small set of prepared cases. Whoever presents picks the files that represent the model best, the environment is clean, and the result appears in seconds. That answers one question: can the model produce the expected output on a chosen example. It is useful evidence about the technology, but does not establish how the real process performs.

What the demo leaves out is exactly what decides your build:

  • the quality of the files that actually arrive;
  • the variation between suppliers and formats;
  • the effort to integrate with the ERP or the database;
  • the daily volume;
  • the cost of an error;
  • and who reviews what the machine cannot solve.

In the document example, a demo shows one clean PDF read correctly. Your operation asks what happens with a crooked scan, an empty field, a value that does not match the order. The demo does not answer that, and applause does not answer it either.

The criteria: what the pilot must prove

A pilot tests a bounded workflow with data, users and criteria set before the build. For document extraction, the questions become numbers the business sets, not the technology:

  • which fields must be extracted and at what acceptable level of accuracy, per field;
  • what volume of documents the flow must process in a day;
  • how long the process takes today and how long it should take;
  • and how mismatches are routed to review.

An acceptable accuracy depends on the cost of an error in that field. A field that decides a payment has different requirements from a field used only for reporting. The diagnostic is usually planned for 2 to 3 weeks. That is when the criteria, data and scope are settled before any build. Compare per-field accuracy, total handling time and review volume against the current process.

Set the criteria first so you cannot move the bar to match the result. Criteria set after the result become a justification, not a measurement.

Demo or pilot: the comparison

Side by side, using the document example, the two options answer different questions and expose different risks:

DimensionDemoPilot
Question answeredCan the model do it on a chosen example?Does the process improve with real data and exceptions?
DataPrepared cases, picked by the presenterReal files from the operation, over a defined period
VolumeA handful of ideal examplesDaily volume, crooked scans included
ExceptionsNot on the agendaRouted to a named review queue
Cost of being wrongA project built on an untested assumptionOne measured period, with criteria to stop

The pilot is usually estimated at 4 to 6 weeks after the scope is defined, subject to access, data and owners being available. That buys the answer the demo cannot give.

When building right after the demo can be right

Building directly can make sense if:

  • the flow is small, cheap to build and fully reversible;
  • a person checks the output before an error can affect the process;
  • and the volume is low enough that manual review covers everything.

Notice these are conditions about consequence, not about the model. When the worst case is cheap, the evidence bar drops. When a wrong output releases a payment, touches a contract or feeds a compliance record, require controlled validation before increasing the investment.

Scale, adjust, or stop

In the document example, a blank field or a value that disagrees with the purchase order goes to review. If scanned files cannot yield the required fields, adjust the data and test again; if the agreed criteria still cannot be met, stop before expanding the flow.

At the end of the pilot, the observed results are compared with the criteria set at the start. The comparison produces a decision with three exits, and all three are legitimate:

flowchart TD
  A["Criteria set before the build"] --> B["Run the pilot"]
  B --> C["Measure against criteria"]
  C --> D{Result vs criteria?}
  D -- met --> E["Scale"]
  D -- retest justified --> F["Adjust scope or data"]
  D -- missed --> G["Stop"]

Scale when the observed result meets the criteria and the dependencies, such as access, data and users, are resolved. Adjust when criteria are missed but a documented change to scope or data justifies another test. Stop when the process did not prove suitable, the data cannot support the flow or the benefit does not appear. Stopping early costs less than scaling in the dark, and a pilot framed as a question makes that exit possible without drama.

In summary

The demo showed a result on a chosen example; the pilot tests the process with real data and exceptions. The decision is whether your process improves with it, and that is a question a pilot answers with numbers set in advance: fields, accuracy per field, volume, time and exception routing. Build straight after the demo only when a wrong output is cheap and reversible. Otherwise, run the pilot, measure against the criteria, and choose between scaling, adjusting or stopping.

Related reading: How to identify where AI makes sense in your operation.

Want to define what a pilot must prove in your operation? Request a diagnostic