← Back to blog

One Page Pilot Checklist: Success Criteria for Project Managers

September 8, 2026
One Page Pilot Checklist: Success Criteria for Project Managers

Pilot success criteria are four pre-committed elements: a metric, a numeric threshold, a measurement method, and a named decision owner. Write these down before launch, on one page, signed and dated. Skip that step and the pilot's outcome gets decided by momentum, calendar pressure, or whoever argues loudest in the final meeting rather than by evidence.


TL;DR:

  • Clear baseline measurement and realistic uplift targets must be established before the pilot begins, based on current performance data.
  • Only a few key metrics, such as adoption rate or error rate, should be tracked within the pilot, with precise thresholds for success, iteration, and failure.
  • A single decision owner with explicit authority should review progress regularly, using pre-signed criteria to determine whether to scale, re-run, or stop the pilot.
  • Automated system logs are the most reliable measurement source, complemented by surveys and observation when necessary, with methods precisely documented.
  • AI-enabled pilots require additional planning, including data contracts, action boundaries, and human review rules, to satisfy regulatory and audit standards.

Aithea
Make Technology Decisions With Confidence
Aithea helps compliance teams navigate technology vendors, prepare RFPs, and manage procurement in regulated financial crime compliance.
Explore Aithea

Table of Contents

What are pilot success criteria and why do they matter?

Pilot success criteria are the pre-agreed definition of what winning looks like: which metric counts, what number it must hit, how that number gets measured, and who has authority to call the result. Without all four, a pilot has no way to fail on the record, only to drift.

Drift is the real enemy of a good pilot. It shows up as a calendar-driven rollout ("the pilot's been running three months, let's just go live"), consensus-driven rollout ("everyone in the room seemed happy"), or politics-driven rollout ("the sponsor already told the board it worked"). None of those are evidence. A government pilot-to-production checklist treats validation as a distinct, scheduled step precisely because it doesn't happen by accident.

Signed criteria matter because they remove the option to rewrite the target after seeing the result. That single habit is what separates a pilot with genuine evidentiary value from an expensive demo:

  • A metric everyone agrees measures the actual business question, not a proxy for it
  • A numeric threshold set before anyone has seen the data
  • A measurement method specific enough that two different people would get the same number
  • A named decision owner with actual authority to stop or scale the pilot

Building the pilot decision statement

Every pilot needs a single sentence that survives the final meeting intact. It has three parts: the decision date, the decision itself (scale, iterate, or stop), and the basis for that decision. Something like: "On 15 May 2026, we will scale this pilot to all three regional teams if adoption exceeds 70% and manual review time drops by at least 20%, measured via system logs." That sentence, agreed before launch, is what stops a pilot report turning into a negotiation.

Build the rest of the framework around it in this order:

  1. Pick one or two headline metrics. Resist the temptation to track fifteen things. A pilot window is short; spread attention too thin and nothing gets measured well enough to trust.
  2. Set usage milestones inside the window. These are checkpoints, not the final verdict, for example "40% of target users active by week two."
  3. Schedule stakeholder reviews at fixed points, not "whenever it feels ready."
  4. Define the expansion trigger in the same numeric terms as the headline metric, so a green result has an obvious next step.
  5. Define the failure criteria with equal precision. If you can't write down what failure looks like, you haven't actually defined success.

Guides aimed at B2B pilot design push this further by tying the decision sentence to a commercial outcome such as paid conversion, contract expansion, or a clear no, rather than a vague "the client liked it," which gives the pilot decision statement a commercial edge as well as an operational one.

Pro Tip: Write the failure criteria before you write the success criteria. Teams find it far easier to spot self-serving thresholds when they've already committed to what "no" looks like.

Building the pilot decision statement — overview diagram

Choosing KPIs: leading and lagging indicators

A good pilot KPI has three qualities: it moves during the pilot window, it can actually be measured with the tools you have, and it connects directly to the decision sentence rather than sitting there as a nice-to-know. Anything else is a vanity metric dressed up as evidence.

Leading indicators tell you what's happening now, inside the pilot, before the bigger business outcome has had time to show up:

  • Adoption rate: the share of the target group actually using the pilot in a given week
  • Error rate: how often the process, tool, or AI system produces an incorrect or flagged result
  • Time-to-complete: whether a task genuinely takes less time than the process it's replacing
  • Escalation frequency: how often a human has to step in and override the pilot workflow

Lagging indicators arrive later and prove the business case, but a short pilot window often can't wait for them:

  • Cost savings, once enough cycles have run to trust the average
  • Contract signatures or renewals, in commercial pilots
  • Retention or churn shifts, which usually need months rather than weeks

Both indicator types earn their place when the numbers are real rather than illustrative. An evaluation of a guaranteed basic income pilot in St Louis found an average credit score increase of about 12 points, measured through administrative records rather than participant self-reporting, which is exactly the kind of objective, hard-to-fake data source a pilot's measurement method should aim for. A separate grocery delivery pilot in Cincinnati recorded a 67% increase in households reporting balanced meals, a fast-moving leading indicator that gave the programme evidence long before any lagging health outcome could appear.

How do you set thresholds and kill or scale rules?

Set the threshold from a real baseline, not a hopeful one. Measure current performance for two to four weeks before the pilot starts, then define the uplift you'd actually expect to see given the pilot's scope and duration, not the uplift that would make the business case look best.

Codify what happens at each result:

  • Green (scale): metric clears the pre-agreed threshold by the review date
  • Yellow (iterate): metric is close but short, triggering a defined re-run window rather than an automatic pass
  • Red (stop): metric misses by a wide margin, or a precondition (data availability, stakeholder engagement) never materialised

Write the rule in numbers, not adjectives: "Scale if adoption exceeds 65% by day 30; re-run with adjusted scope if adoption is 45 to 65%; stop below 45%." A rule that says "scale if adoption is strong" gives everyone room to argue after the fact.

Pro Tip: Build the yellow category deliberately. Most pilots that get killed unfairly are actually yellow results being forced into a red/green binary because no one defined a middle path in advance.

Who owns the go/no-go decision?

Name one decision owner, not a committee, and give them explicit authority to stop the pilot even if the sponsor wants it to continue. Diffuse ownership is how weak pilots survive past the point they should have ended.

Set three checkpoints as a minimum cadence:

  • Weekly check-in: data quality and early usage signals only, no verdict discussion
  • Midpoint review: compare actuals against the milestone targets set in the decision statement
  • Final review: the decision owner presents the scored result against the pre-agreed thresholds

Map the stakeholder matrix explicitly: the economic buyer who controls budget for scaling, the daily users whose adoption data you're measuring, the technical owner responsible for system logs and data integrity, and the executive sponsor who needs the one-page summary rather than the full appendix. The Executive FX checklist for CFOs offers a useful parallel: a fixed short cadence of review points, built to stop decisions sliding by default rather than by design.

Which measurement methods should you trust?

System logs are the most reliable source when the pilot runs inside software you control. They're automatic, hard to game, and timestamped, but they only capture what the system was built to record.

Surveys fill the gap for anything system logs can't see, satisfaction, perceived friction, willingness to recommend, but they're slower to collect and prone to social pressure bias, particularly if the person running the survey also runs the pilot.

Observation (someone watching the workflow in real time) catches issues neither of the above will flag, especially workarounds staff invent to make a clunky pilot tool usable, but it doesn't scale past a handful of sessions.

Document the method precisely enough that two different people would measure the same thing the same way: which system, which field, which time window, which survey question wording. When a single source feels incomplete, triangulate: cross-check system log adoption figures against a short survey asking why people who didn't log in stayed away.

Comparison of three pilot measurement methods

What are the early warning signs a pilot is failing?

Four signals show up early and reliably: adoption stalling well below the milestone target, gaps or inconsistencies in the data you need to score the pilot, the economic buyer going quiet on reviews, and the scope quietly expanding beyond what the decision statement covers.

Each has a proportionate response rather than an automatic kill:

  • Low adoption: narrow the pilot to the strongest use case rather than abandoning it outright
  • Missing data: convert the remaining weeks into a learning pilot focused on why the data gap exists
  • Absent decision owner engagement: escalate once, in writing, before the final review, not after it
  • Scope creep: cut back to the original decision sentence; anything added mid-pilot doesn't count towards the original threshold

Write failure reasons specifically. "Adoption failed because the tool required a login step most users skipped" preserves organisational learning; "the pilot didn't work" preserves nothing.

How do you score and present the final result?

Score five categories rather than a single number, then let the combination point to a recommendation:

  1. Business outcome: did the headline metric clear its threshold?
  2. Usage: did adoption meet the milestone targets across the pilot window?
  3. Stakeholder confidence: did the buyer, users, and technical owner all stay engaged through to the final review?
  4. Implementation effort: was the pilot's operational burden sustainable at scale?
  5. Commercial path: is there a credible budget or contract route to full deployment?

Score each green, yellow, or red. Mostly green maps to scale; a mix with one or two yellows maps to iterate with a defined re-run window; more than one red maps to stop. Academic work on composite scoring in pilot evaluation supports combining several subscores into one defensible rating rather than relying on a single metric that might mask a weak result elsewhere.

Present it as one page for executives, scores and the recommendation only, with a full evidence appendix behind it for anyone who wants to check the working.

The one-page pilot success criteria checklist

Copy this into your pilot charter and fill each line before launch:

  • Decision sentence: date, decision options, and the basis for choosing between them
  • Timeline: start date, review dates, final decision date
  • Scope: exactly which users, workflow, or process the pilot covers, and what's explicitly excluded
  • Primary outcome: the one or two headline metrics
  • Thresholds: the numeric green/yellow/red boundaries for each metric
  • Measurement method: system, field, and time window for each metric
  • Usage milestones: interim checkpoints inside the pilot window
  • Expansion trigger: what a green result unlocks, in concrete terms
  • Failure criteria: what a red result looks like, written before the pilot starts
  • Decision owner: named individual, with sign-off

Adapt the scope line for the pilot type: an AI agent pilot needs a data contract and an action boundary added; a process-change pilot needs the specific teams and volumes named; a product trial needs the customer segment defined. Keep the whole artefact to one page, and have the decision owner sign it before day one.

Guidance for regulated and AI-enabled pilots

Compliance pilots running across European jurisdictions carry an extra constraint: measurement methods must be built for audit from the outset, with data minimisation and retention rules baked into the design rather than retrofitted once a regulator asks. For AI agent pilots specifically, add a data contract, a defined action boundary, and human review rules before launch, alongside the standard exit criteria covering ship, narrow, extend, or stop. A structured technology selection process helps teams choose vendors whose systems can actually produce the audit trail a regulated pilot needs.

Three rules that keep pilots honest

Lock the decision sentence before launch and don't touch it once data starts arriving. Require a measurement method independent of whoever's championing the pilot, ideally system logs rather than self-reported wins. Schedule the decision meeting on day one and enforce quorum. Most pilot drift isn't dishonesty, it's a sponsor who's already emotionally invested. Naming that pressure early, in the room, before results exist, is usually enough to defuse it.

— Aneta

How AITHEA supports pilot design and technology selection

Designing a defensible pilot gets harder once regulated data, cross-border rules, or AI agents enter the picture, which is exactly where most teams need outside eyes before they need outside tools. Specialist firms work with compliance and risk teams on pilot design, vendor shortlisting, and procurement for financial crime compliance technology, including AI-enabled workflows that need a data contract and human review rules built in from day one.

Aithea

When the pilot involves an AI agent, cross-border data, or a regulated workflow, the design questions multiply fast: what counts as an auditable log, who reviews an AI decision, what happens when a vendor's data handling doesn't match your retention obligations. Aithea's Heliolus technology selection navigator exists to help compliance teams work through exactly that shortlist before a contract gets signed, and the consultancy team can help draft the pilot decision statement itself. If your next pilot touches AI, sanctions data, or KYC workflows, consider seeking external expertise to scope the pilot design before you set a launch date.

Sources

FAQ

What are examples of pilot success metrics?

Common examples include adoption rate, error rate, time-to-complete a task, escalation frequency, and cost savings, chosen so at least one metric is measurable inside the pilot window itself.

How many success criteria should a pilot have?

Keep to one or two headline metrics with clear thresholds. Tracking too many during a short pilot window spreads attention thin and weakens the evidence behind any single number.

Who should be the decision owner for a pilot?

One named individual with real authority to stop or scale the pilot, typically the economic buyer or executive sponsor, not a committee that can dilute accountability at the final review.

What causes most pilots to fail?

Low adoption, missing or inconsistent measurement data, an absent or disengaged decision owner, and scope creep are the most common early warning signs of pilot failure.

How does AI change pilot success criteria?

AI agent pilots need extra elements beyond a standard metric and threshold: a defined data contract, an action boundary limiting what the system can do unsupervised, and human review rules, all agreed before launch.