Skip to main content
Download free report
Softblues
Softblues
Back to Blog
AI Strategy & Consulting
July 21, 20268 min read

How to Run an AI Proof of Concept That Reaches Production: A UK Guide

Gartner expected up to 30% of generative AI projects to be abandoned after the proof of concept. Most stall because the PoC proved the wrong thing. Here is a fixed six-week plan to run one that reaches production.

How to Run an AI Proof of Concept That Reaches Production: A UK Guide

By Ivan Pylypchuk, CEO of SoftBlues. Has led Claude implementations from proof of concept to production for finance, legal and healthcare teams across the UK and Ireland.

Gartner predicted that at least 30% of generative AI projects would be abandoned after the proof of concept by the end of 2025, blaming poor data quality, weak controls, rising costs and unclear business value (Gartner, July 2024). Most of those projects were not killed by the technology. They stalled because the proof of concept proved the wrong thing.

A proof of concept (PoC) is a small, time-boxed build that tests whether an AI approach can solve a specific business problem well enough to justify going further. Done well, it de-risks the decision to scale. Done badly, it produces an impressive demo that nobody can put into production. At SoftBlues, an AI consulting firm working with regulated mid-market companies across the UK and Ireland, we run PoCs on a fixed six-week plan designed to answer one question: should we build this for real? This is that plan.

Key facts

  • A proof of concept tests one business problem against real data on a fixed timeline, to decide whether to build for production, not to produce a demo.
  • Gartner expected at least 30% of generative AI projects to be abandoned after PoC by end of 2025 (Gartner, July 2024).
  • The failures cluster on unclear value, poor data and weak controls, all decidable before you build rather than after.
  • A good PoC defines its success metric and its kill criteria on day one, uses your real data, and produces a decision, not just a demo.
  • Our PoCs run to a fixed six-week plan: define, access data, build, test against the metric, review, then decide to scale, adjust or stop.
  • Who this is for, and who it isn't

    This is for a decision-maker at a 50–500-person UK or Ireland company who has an AI idea and a budget to test it, and wants the PoC to end in a clear go or no-go rather than an open-ended experiment.

    It is not for a team that already has a validated use case and needs a production build, and it is not a deep technical tutorial. If you are still choosing between partners, read our guide to what should be in an AI consulting proposal first.

    Why do so many AI proofs of concept fail to reach production?

    They fail for reasons you can see before you start. Gartner's own breakdown points to data quality, cost, controls and unclear value rather than model capability. In our experience the single most common cause is that the PoC had no agreed definition of success, so "it works" became a matter of opinion and the project drifted.

    The second cause is demo data. A PoC built on a clean, curated sample looks brilliant and then collapses on the real inputs, because production data is messier than anyone admits. The third is that the PoC ignored the boring parts (access, permissions, integration, governance) that decide whether a working model can actually be deployed in a regulated business.

    Warning
    A demo proves the model can do the task once, on good data, in a controlled setting. Production needs it to do the task reliably, on your data, inside your controls. A PoC that only produces a demo has not reduced your real risk.

    What makes a proof of concept worth running?

    A PoC is worth running when there is a real decision waiting on the result and you have defined what would make the answer yes. Before any build, pin down four things: the specific problem, the metric that defines success, the data you will test on, and the criteria that would make you stop.

    The problem should be narrow. "Use AI in finance" is not a PoC; "extract line items from supplier invoices and match them to purchase orders with 95% accuracy" is. The metric should be measurable and agreed with the people who will own the result. The data must be your real data, not a curated sample. And the kill criteria matter as much as the success metric: knowing in advance what result would make you walk away is what turns a PoC into a decision rather than a sunk cost.

    How to run an AI proof of concept: a six-week plan

    Here is the plan we use. It is deliberately short. A fixed timeline forces the scope to stay narrow and the decision to stay real.

    Week 1: Define. Write down the one problem, the success metric with a number, the kill criteria, and who signs off the go or no-go. Agree the scope in writing so it cannot creep.

    Week 2: Get real data and access. Pull a representative sample of your actual data, including the awkward edge cases. Sort out permissions and access early, because in regulated businesses this is what quietly kills projects later.

    Weeks 3 and 4: Build the thin slice. Build only enough to test the metric end to end on real data. No polish, no scope creep, no nice-to-haves. The goal is a working path, not a product.

    Week 5: Test against the metric. Run the build on the full sample and measure against the number you set in week one. Record where it fails and why, not only the headline score.

    Week 6: Review and decide. Compare the result to the success metric and the kill criteria, cost the production build, and make an honest go, adjust, or stop call. The output of week six is a decision with evidence behind it.

    💡Tip
    Set the kill criteria on day one, when nobody is emotionally invested. It is the cheapest insurance you can buy against a project that limps on because stopping feels like failure.

    Proof of concept, pilot or production: what is the difference?

    These words get used interchangeably and it causes confusion in scoping. They are three different stages with three different goals.

    StageQuestion it answersScopeRuns on
    Proof of conceptCan this work well enough to justify building it?One narrow problem, thin sliceA real data sample
    PilotDoes it hold up with real users in a limited rollout?One team or workflowLive use, limited scope
    ProductionCan it run reliably for everyone, under our controls?The full processLive, at scale, governed

    A PoC that quietly turns into a permanent pilot with no decision point is one of the patterns Gartner is describing. Each stage should end with an explicit decision to move to the next.

    What does an AI PoC cost, and what should you get for it?

    The cost depends on the problem, but the deliverables should be fixed regardless. A PoC should hand you a working thin-slice build, a measured result against the agreed metric, the failure cases it found, and a costed recommendation for production. If a partner cannot tell you the deliverables and the timeline up front, that is a warning sign.

    (Our engagement model, not a published benchmark: we run PoCs on a fixed six-week plan and put validated use cases into production in 90 days at a fixed price, with a money-back guarantee if it fails to deliver the agreed outcome. Figures are indicative, 2026.)

    For a worked example of a PoC that led to a rebuild rather than a rushed rollout, see our anonymised Claude Code audit case study, where the proof-of-concept stage surfaced the risks before anything went live. If your end goal is running the business on Claude, our Claude operating system case study shows what production looks like after the PoC.

    What are the red flags in an AI PoC?

    1. No success metric. If nobody has written down the number that means "yes", the PoC cannot end in a clear decision.

    2. Demo data instead of your data. A PoC on a clean sample is a sales demo. Insist on your real, messy inputs.

    3. No kill criteria. Without an agreed stopping point, a failing PoC drifts into a permanent, expensive pilot.

    4. Ignoring access and governance. In FCA, SRA or CQC-regulated work, data access, permissions and controls decide whether a working model can ever be deployed. A PoC that skips them has not tested the real constraint.

    Questions to ask before you commission a PoC

    1. "What is the success metric, and who agrees it?" You want a number and a named owner, not "we'll see how it goes".

    2. "What data will you test on?" The right answer is your real data, including edge cases, not a curated sample.

    3. "What are the kill criteria?" A serious partner will help you define when to stop before you start.

    4. "What do I get at the end, and by when?" Look for fixed deliverables and a fixed date, ending in a costed go or no-go.

    If you want to understand the road from here, read our AI implementation roadmap on moving from pilot to production and our framework for measuring AI ROI.

    Frequently asked questions

    How long should an AI proof of concept take?

    Long enough to test the metric on real data, and no longer. We run PoCs on a fixed six-week plan. A tight timeline keeps the scope narrow and forces a real decision at the end rather than an open-ended experiment.

    What is the difference between a PoC and a pilot?

    A proof of concept tests whether an approach can work, on a real data sample, to decide whether to build. A pilot tests whether the built solution holds up with real users in a limited rollout. The PoC comes first and should end with a go or no-go.

    Why do so many AI PoCs fail?

    Gartner attributes the failures to poor data quality, weak controls, rising costs and unclear business value rather than model capability. In practice the biggest avoidable cause is running a PoC with no agreed success metric or kill criteria.

    Should a PoC use real data or sample data?

    Real data. A PoC built on a clean, curated sample overstates how well the approach will work, because production inputs are messier. Testing on your real data, including the awkward cases, is the point.

    What should I get at the end of a PoC?

    A working thin-slice build, a measured result against the agreed metric, the failure cases it found, and a costed recommendation for whether and how to move to production. In short, a decision with evidence, not just a demo.

    Can a PoC be run on regulated data?

    Yes, with the right controls: agreed data access, permissions, and a human reviewing anything consequential. For FCA, SRA or CQC-regulated work, governance is part of what the PoC has to prove, not an afterthought.

    What happens if the PoC fails its metric?

    That is a successful PoC. It has told you, cheaply and quickly, not to spend production money on the wrong approach. You either adjust the scope and retest or stop. Both outcomes are worth far more than a demo that hides the risk.


    SoftBlues is a registered Anthropic Partner Network member and a Google Cloud Partner. We are practitioners, not slide-deck consultants: we run PoCs to a fixed six-week plan that ends in an honest go or no-go, and we put validated use cases into production in 90 days at a fixed price, with a money-back guarantee if it fails. If you have an AI idea worth testing properly, book a discovery call.

    See it in production

    Systems we have built and run for clients, with the numbers that came out of them.

    Browse all case studies

    Related Articles