The AI proof of concept has a familiar pattern. A small team takes a model, a sample of data and a controlled task. The demonstration works. Stakeholders are impressed. The project is described as successful.
Then progress slows.
The data is not available in the real workflow. Users need to review every result. Exceptions create more work than expected. The system has no owner. Accuracy is measured, but the operational or financial consequence is not. A production estimate arrives and the value case no longer holds.
None of this means the proof of concept failed. It means it proved too little.
What a proof of concept proves
A conventional proof of concept is designed to answer a technical question such as:
- Can the model classify these documents?
- Can an agent complete this sequence?
- Can retrieval produce relevant answers from this knowledge base?
- Can a forecast outperform the current method on historical data?
These are legitimate questions. They reduce technical uncertainty.
But the investment decision usually depends on a wider system. A model can perform well while the end-to-end capability creates no material value.
What a Minimum Viable Capability proves
A Minimum Viable Capability, or MVC, is the smallest credible version of the business capability required to test the value logic in real conditions.
It includes only what is necessary across six elements:
- Technology: the model, agent, automation or product needed for the consequential task.
- Human judgement: the decisions people retain, the review burden and the accountability model.
- Workflow: where the capability enters the work, what it replaces and how exceptions flow.
- Data: representative inputs, quality requirements and feedback signals.
- Controls: the safeguards appropriate to the consequence of failure.
- Measures: the baseline and indicators that connect performance to value.
The MVC is not a miniature production platform. It is a deliberately incomplete capability built to answer the questions that determine whether production investment is justified.
Proof of concept and Minimum Viable Capability compared
| AI proof of concept | Minimum Viable Capability | |
|---|---|---|
| Primary question | Can the technology do the task? | Can this capability change the outcome? |
| Environment | Controlled | Representative real work |
| User involvement | Often demonstration or feedback | Active participation in the workflow |
| Human judgement | Commonly outside scope | Designed explicitly |
| Baseline | Technical benchmark | Operational and economic baseline |
| Measures | Accuracy, latency, task completion | Quality, time, effort, capacity, value and risk |
| End decision | Technically feasible or not | Scale, revise, narrow, defer or stop |
How to scope an MVC
Start with the invalidating assumptions
Do not begin with a feature list. Begin with the assumptions that could make the value case fail.
For example:
- users may not trust the output enough to act;
- review effort may consume the time saved;
- the relevant data may arrive too late;
- the capability may improve speed but reduce conversion or quality;
- the organisation may be unable to remove the old work;
- production cost may exceed the value created.
The MVC should be designed to expose those risks quickly.
Choose representative work, not convenient work
Clean historical examples make demonstrations easier. They can also hide the very exceptions and operating conditions that determine whether the capability is useful.
Select a bounded cohort of real or faithfully representative work. Include enough variation to learn without expanding into a full rollout.
Make the human role visible
“Human in the loop” is too vague to be a design.
Specify:
- what the person sees;
- what decision they make;
- when they must intervene;
- how much time review takes;
- what expertise is required;
- who is accountable for the result;
- what feedback improves the capability.
If every output needs expert rework, the economics may be worse than the model metrics suggest.
Define the threshold before the test
Agree what result would justify the next investment before enthusiasm or sunk cost affects the decision.
Thresholds might include:
- acceptable quality or error rates;
- reduction in handling time or delay;
- conversion or revenue effect;
- capacity released and whether it can be redeployed or removed;
- user adoption and exception burden;
- cost per completed outcome;
- risk and control performance.
Keep integration proportional to the question
Some integration may be essential to create a credible test. Full production architecture rarely is.
Use the least integration needed to test real workflow behaviour. A manual bridge is acceptable when it does not hide a material cost or risk. Document what would change at scale.
What a successful MVC produces
The most valuable output is not the prototype. It is the improved quality of the investment decision.
A successful cycle produces:
- observed performance in representative conditions;
- a clearer value model;
- evidence about user behaviour and review effort;
- known data, control and integration requirements;
- a view of the workflow and role changes required;
- a scale, revise or stop recommendation.
The result may be to stop. That is still value if it prevents a larger investment in a weak opportunity.
The practical standard
Do not ask only whether the AI works.
Ask whether the smallest credible business capability changes the work enough to create the outcome—and whether the evidence has earned further investment.