Write the exit question before building the pilot.
State the business outcome, current baseline, target users, required behaviour, acceptable risk, human-control model, production assumptions, and evidence leadership needs for the next decision.
A useful pilot finishes with evidence to proceed, change the design, prepare missing foundations, pause, or stop.

ONE BRIEFING · 6 DECISION LAYERSState the business outcome, current baseline, target users, required behaviour, acceptable risk, human-control model, production assumptions, and evidence leadership needs for the next decision.
Test a complete journey with representative data, real rules, actual or production-like integrations, appropriate permissions, named users, ordinary exceptions, and the expected handoffs.
Test missing information, weak evidence, unavailable systems, duplicate events, policy conflicts, low confidence, denied access, rejected approvals, intervention, rollback, and recovery.
Proceed, refine, redesign, prepare foundations, or stop. Record the evidence, unresolved risk, architecture, controls, resources, ownership, operating cost, and next-stage scope behind that answer.
A production decision must include integration hardening, identity, monitoring, support, incident response, documentation, training, source maintenance, release control, cost management, and the responsible operating team.
A successful controlled launch does not automatically justify more users, markets, languages, workflows, or autonomy. Each expansion changes the operating and risk profile and deserves its own decision.
A controlled pilot should expose the important production assumptions without putting the full organisation at risk. Before building, leadership should know which decision the pilot will inform, which evidence is required, who will decide, what would prevent progress, and what additional work a successful result would still require before production.
Define the intended business outcome, baseline, users, process, required behaviour, acceptable risk, operating cost assumptions, and production question. The exit should allow several honest answers: proceed, refine, redesign, prepare foundations, or stop. Agree who decides and which evidence matters before enthusiasm about the prototype changes the standard.
Use representative people, data, rules, permissions, integrations, ordinary exceptions, and human handoffs. A disconnected sandbox using perfect examples may prove that a model can produce an output but not that the business can use it. Protect sensitive environments appropriately, but preserve the constraints most likely to determine reliability, adoption, security, and operating effort.
Test missing or conflicting information, denied access, unavailable tools, duplicates, delays, weak evidence, policy conflict, user correction, escalation, rollback, and recovery. Measure the final outcome, source support, tool behaviour, human effort, waiting, errors, overrides, cost, and user experience. Record limitations and failures rather than tuning the demonstration until only successful cases remain visible.
A positive pilot does not include hardened integrations, identity, capacity, monitoring, support, incident response, documentation, training, source maintenance, release control, supplier management, or continuous evaluation by default. Estimate these requirements and assign owners before approving production. Treat more users, markets, languages, workflows, data, and autonomy as separate evidence gates because each changes the risk and operation.
Agree who decides, what evidence they need, which thresholds matter, and which risks would prevent production before the pilot begins.
Use representative people, data, rules, permissions, integrations, exceptions, and handoffs while keeping volume and audience controlled.
Include monitoring, review, support, maintenance, vendor usage, evaluation, release, and ongoing improvement in the production decision.