Skip to main content
Download free report
Softblues
Softblues
Back to Blog
AI Strategy & Consulting
July 31, 202610 min read

How to Choose an AI Assistant Development Team: 10 Questions That Separate Builders from Talkers

MIT found 95% of generative AI pilots produce no measurable return. Almost none of that is model quality. Here are the ten questions to put to an AI assistant development team before you sign, and the answers a team that ships actually gives.

How to Choose an AI Assistant Development Team: 10 Questions That Separate Builders from Talkers

By Ivan Pylypchuk, CEO of SoftBlues

MIT's Project NANDA looked at 300 public enterprise AI deployments and found that 95% of generative AI pilots produced no measurable return (MIT Project NANDA, July 2025). Almost none of that failure is model quality. It is the team you picked, and the questions you did not ask them before you signed.

An AI assistant is not hard to demo. Anyone can wire a chat window to a language model in an afternoon and show you something impressive on a call. Getting that same assistant to answer correctly against your real documents, with your real permissions, for a year, with someone accountable when it is wrong, is a different job. This post is the vetting list we would want a buyer to use on us.

Key facts

  • 95% of generative AI pilots delivered no measurable P&L impact (MIT Project NANDA, 300 deployments reviewed, July 2025).
  • The report attributes the failure to brittle workflows and weak integration with day-to-day operations, not to model quality.
  • UK market daily rates for a senior AI engineer sit between £600 and £1,200; a middle AI engineer between £450 and £850 (SoftBlues rate card, UK market comparison, 2026).
  • A typical two-month proof of concept for an AI assistant runs at around £20,000 with one to two engineers plus part-time senior and PM support (SoftBlues rate card, PoC phase).
  • Audience: UK and Ireland companies with 50 or more knowledge workers.
  • What is an AI assistant development team, and who is actually on it?

    An AI assistant development team is the small group that designs, builds, evaluates and runs a language-model application against your own data and systems. Four roles do the work. An AI engineer builds the retrieval layer, the prompts and the tool calls. A full stack engineer builds the interface and the integrations into your systems. A QA engineer builds the evaluation set, which is the part most teams skip. A project manager keeps scope and access moving, because most AI projects stall on someone else's IT permissions rather than on code.

    Team size matters less than composition. A three-person team with a named QA engineer and a written evaluation set will beat a six-person team without one, because an assistant that cannot be measured cannot be improved. Ask who writes the evaluation set and how many test questions it holds. If nobody owns that, you are buying a demo.

    Four cards showing the things to vet in an AI assistant development team: sector proof, data handling, a named team, and a route to production.

    What should you ask before you sign?

    Ten questions. Each one has an answer a team that ships gives, and an answer that should worry you.

    1. Show me an assistant you built that is still running. A good team names the system, the year it went live, roughly how many people use it, and what broke in month three. A weaker answer is a slide of logos with no live system behind them.

    2. Who exactly will do the work, and are they available now? Ask for names, seniority and start dates. Selling teams often pitch with senior people and staff with juniors. Put the named people in the contract and ask what happens if one leaves mid-project.

    3. Where does my data go, and who can read it? You want a plain answer: which model provider, which region, whether your prompts train anyone's model, who on their side has production access. Vague reassurance here is the single biggest red flag in the whole list.

    4. How will we know if the assistant is right? The answer should be an evaluation set built from your real questions with expected answers, scored before launch and re-scored after every change. "We'll test it as we go" means there is no measurement.

    5. What happens when it does not know? Good assistants refuse, cite the document they used, and hand off to a person. A team that has not thought about refusal behaviour has not run one of these in front of real staff.

    6. What is fixed and what is not? Ask for a fixed price and a fixed scope on the first phase, with the assumptions written down. Open-ended time and materials on a first AI build transfers all the risk to you.

    7. Who owns the code, the prompts and the evaluation set? All three should be yours, in your repository, from day one. Prompts and evaluation sets are the real intellectual property in an assistant, and they are the easiest thing for a supplier to keep hold of.

    8. What does it cost to run every month after launch? Model usage, hosting, monitoring and someone's time. A team that cannot give you a monthly running cost has not operated one of these at scale.

    9. What would you tell me not to build? Any team worth hiring has a list. If everything you describe sounds like a great fit for their services, you are talking to a sales function.

    10. What does the handover look like? Documentation, a runbook, a named owner on your side, and a support arrangement with response times. Ask to see a handover pack from a previous project with the client details removed.

    Important
    Questions 3 and 4 are the ones that predict outcomes. Data handling tells you whether the project will survive your own security review. The evaluation set tells you whether anyone will be able to prove the assistant works six months from now.

    Want a second opinion on a team you are already talking to? A 30-minute discovery call costs nothing and is not a pitch. Bring their proposal and we will tell you which parts we would push back on, where the scope is likely to move, and what we would build differently. You leave with a shortlist of questions to put to them. Book a discovery call.

    How do you tell a builder from a talker?

    The difference shows up in what they bring to the second meeting. Builders arrive with a narrowed scope, a rough evaluation plan and a question about your data access. Talkers arrive with a longer deck.

    Two-column comparison contrasting builders, who show shipped systems, named engineers and a fixed scope, with talkers, who bring slide decks, an anonymous bench and open-ended scope.

    SignalBuilderTalker
    ProofA running system you can look at, with its limits describedLogos, awards, headcount
    TeamNamed engineers, CVs, availability dates"Our expert team"
    ScopeOne process, fixed price, written assumptionsA phased transformation programme
    DataAsks about permissions and retention in meeting oneRaises security at contract stage
    MeasurementOffers an evaluation set and a pass markPromises accuracy as an adjective
    HonestyTells you what not to buildEverything is a fit
    Best forCompanies that want one process working in productionCompanies that need a board-level strategy artefact
    Avoid ifYou genuinely need a market-wide strategy review firstYou need something live this quarter

    We built SofiaHR's candidate assessment platform this way. The useful detail for a buyer is not the technology. It is that the work started by testing ten different models against real assessment data before a line of product code was written, because the choice of model was a measurable question rather than a preference.

    What should an AI assistant actually cost?

    Two numbers to hold in your head. UK market daily rates for a senior AI engineer run from £600 to £1,200, and a middle AI engineer from £450 to £850 (SoftBlues rate card, UK market comparison). A first proof of concept for an assistant, two months, one to two engineers with part-time senior and project management support, prices at around £20,000.

    That range is the honest one for a first phase with a defined scope. Anything materially cheaper is usually a demo with no evaluation layer and no integration work. Anything materially more expensive on a first phase is usually a discovery exercise wearing a build budget.

    PhaseTypical durationTypical teamWhat you should get
    Proof of concept2 months1 to 2 engineers, part-time senior and PMWorking prototype on your data, evaluation set, a written recommendation
    Production build3 to 6 months2 to 4 engineers, QA, PMIntegrated assistant, monitoring, runbook, handover
    Run and improveOngoingFractional engineer plus supportMonthly evaluation scores, model updates, a named owner

    For the fuller picture on budgets and rate structures, we set it out in AI consulting costs in the UK. If you are earlier than that and want to get the first phase right, how to run an AI proof of concept that reaches production covers the phase this post prices.

    Warning
    A fixed price with no written assumptions is not a fixed price. Ask which assumptions, if broken, trigger a change request. The usual ones are data quality, access timelines and the number of source systems.

    What are the red flags?

    Six things that reliably predict trouble.

    No named engineers in the proposal. You are buying capacity from a bench, and you will get whoever is free.

    No evaluation plan. Without one, "it works" is a matter of opinion, and opinions change when the person who liked the demo leaves.

    Accuracy claimed as a percentage with no test set attached. A number with no denominator is marketing.

    A refusal to discuss what happens after go-live. Assistants drift as models update and documents change. A team with no run phase has not lived with one.

    Discovery priced as a large fixed fee with no build commitment. Useful discovery ends in a costed build decision, not a report.

    Everything is a fit. The strongest signal in a sales conversation is a supplier telling you which part of your idea to drop.

    Which questions should your own side be ready to answer?

    Vetting runs both ways, and the answers you cannot give are usually what delays the project.

  • Which single process are we improving, and what is it costing us today in hours or errors?
  • Who owns the documents the assistant will read, and can they grant access in the first fortnight?
  • Who signs off that an answer is correct?
  • What does the assistant have to be right about, and where is a wrong answer merely annoying rather than serious?
  • Who on our side will own it after handover, and how much of their week does that take?
  • If you cannot answer the first and the last of those, fix that before you brief anyone. It changes the price and it changes who you should hire.

    Frequently asked questions

    How long does it take to build an AI assistant? A useful first version against your own documents takes about two months. Production, with integrations, monitoring and a handover, usually adds three to six months depending on how many systems it touches.

    Should I hire an in-house AI engineer instead? If you expect a steady pipeline of AI work for the next two years, yes, eventually. For a first assistant, an in-house hire means you learn the mistakes on your own payroll. Most companies get better value bringing in a team that has already made them, then hiring once the pattern is proven.

    What size team do I need? One to two engineers for a proof of concept. Two to four plus QA and a project manager for a production build. If a supplier proposes more than that for a first assistant, ask what each person does in week one.

    Do I need my own data scientists? No. Modern assistants are built on hosted models, so the work is retrieval, integration and evaluation rather than model training. You need engineers who can measure output quality, not researchers.

    How do I check a supplier's claims? Ask to speak to a client whose system is still running, and ask that client what broke and how quickly it was fixed. Also ask the supplier for an anonymised handover pack. Both requests are ordinary, and reluctance to meet them tells you something.

    What if the assistant does not work? Agree the exit before you start. A first phase should end in a written recommendation that includes the option to stop. We put a money-back guarantee on our 90-day production commitment for exactly this reason.

    Can this run on Claude, or do I need something custom? Most business assistants are better built on a hosted model such as Claude than on anything custom. We run six of our own departments this way, which is documented in our own Claude operating system. Where a Microsoft-native shop is already standardised on Copilot, we will say so.


    SoftBlues is a registered Anthropic Partner Network member, and a registered partner with Google Cloud and Microsoft. We build AI assistants for UK and Ireland companies with 50 or more knowledge workers, fixed price, in production in 90 days, money back if it fails. We use what we sell before we sell it, which is why the questions above are the ones we expect to be asked.

    If you are choosing a team and want the shortlist pressure-tested, book a discovery call.

    See it in production

    Systems we have built and run for clients, with the numbers that came out of them.

    Browse all case studies

    Related Articles