Key takeaways
- Write the decision, hypothesis, baseline, representative conditions, and pass-fail thresholds before the POC begins.
- Keep the test small enough to learn, but not so artificial that success says nothing about the real environment.
- Define access, data use, security, human oversight, failure response, retention, and cleanup for the test itself.
- Agree what happens after pass, partial pass, fail, or inconclusive—especially when the honest answer is stop.
01
The demo worked because the difficult parts were removed
The sample was clean, the integrations were mocked, the evaluator knew which outputs to ignore, and the vendor tuned the workflow between runs. The final deck says 92% accuracy. Nobody can explain the denominator, baseline, excluded cases, or what the number means in production.
Write a POC charter with the business decision, riskiest assumption, users, representative data, baseline, scenarios, exclusions, metrics, thresholds, evaluator, timing, cost boundary, safety controls, and exit path. If the team cannot say what result would make it stop, it is running a sales performance, not a test.
Swipe to compare every column
| Outcome | Meaning | Decision |
|---|---|---|
| Pass | Pre-agreed threshold met under valid conditions | Plan a controlled next stage |
| Partial pass | Some value and material unresolved risk | Narrow, remediate, or retest |
| Fail | Critical threshold or safety boundary missed | Stop or redesign |
| Inconclusive | Test cannot support the decision | Fix the method before claiming a result |
02
Test the risky assumption, not every feature
GOV.UK’s alpha guidance recommends prototypes just substantial enough to test the riskiest assumptions and support a decision about moving forward. Apply the same restraint here. A broader POC creates more activity but can dilute the one uncertainty that should decide whether the investment continues.
Use representative volume, edge cases, user roles, latency, language, accessibility needs, integration constraints, and failure modes. Document every intervention by the vendor or internal expert. Assistance is not automatically disqualifying, but hidden assistance makes the result impossible to interpret.
03
Build a safety boundary around the experiment
Use the minimum necessary data, prefer synthetic or de-identified material when it can answer the question, restrict access, define retention, and make deletion verifiable. Keep the prototype away from production decisions when the risk or evidence is not ready. Provide a human override and a path for affected people to report a problem.
For AI systems, NIST AI RMF provides a useful structure: govern the work, map the context, measure risk, and manage it. A POC score should not erase legal, privacy, security, fairness, or operational constraints that the test did not examine.
04
Publish the limits with the result
Report the sample, baseline, metric definition, confidence or uncertainty where appropriate, excluded cases, failed scenarios, manual interventions, environment differences, and unresolved risks. Do not turn a controlled test into a universal marketing claim.
Track time to decision, evaluator effort, threshold changes after launch, invalid runs, safety incidents, cleanup completion, and production assumptions still untested. A failed POC can be valuable when it prevents an expensive rollout. A passed POC is only permission to take the next bounded step.
Primary sources and further reading
Use the source material to validate details against your own context and current platform configuration.
- GOV.UK Service Manual: How the alpha phase works
- GOV.UK Service Manual: Making prototypes
- NIST: AI Risk Management Framework
This field note follows the XenGrowth editorial policy: primary sources where available, visible limitations, material review dates, and no invented first-hand experience.
Stay with the problem



