How to Choose an AI Assistant Development Team: 10 Questions That Separate Builders from Talkers
MIT found 95% of generative AI pilots produce no measurable return. Almost none of that is model quality. Here are the ten questions to put to an AI assistant development team before you sign, and the answers a team that ships actually gives.

By Ivan Pylypchuk, CEO of SoftBlues
MIT's Project NANDA looked at 300 public enterprise AI deployments and found that 95% of generative AI pilots produced no measurable return (MIT Project NANDA, July 2025). Almost none of that failure is model quality. It is the team you picked, and the questions you did not ask them before you signed.
An AI assistant is not hard to demo. Anyone can wire a chat window to a language model in an afternoon and show you something impressive on a call. Getting that same assistant to answer correctly against your real documents, with your real permissions, for a year, with someone accountable when it is wrong, is a different job. This post is the vetting list we would want a buyer to use on us.
Key facts
What is an AI assistant development team, and who is actually on it?
An AI assistant development team is the small group that designs, builds, evaluates and runs a language-model application against your own data and systems. Four roles do the work. An AI engineer builds the retrieval layer, the prompts and the tool calls. A full stack engineer builds the interface and the integrations into your systems. A QA engineer builds the evaluation set, which is the part most teams skip. A project manager keeps scope and access moving, because most AI projects stall on someone else's IT permissions rather than on code.
Team size matters less than composition. A three-person team with a named QA engineer and a written evaluation set will beat a six-person team without one, because an assistant that cannot be measured cannot be improved. Ask who writes the evaluation set and how many test questions it holds. If nobody owns that, you are buying a demo.

What should you ask before you sign?
Ten questions. Each one has an answer a team that ships gives, and an answer that should worry you.
1. Show me an assistant you built that is still running. A good team names the system, the year it went live, roughly how many people use it, and what broke in month three. A weaker answer is a slide of logos with no live system behind them.
2. Who exactly will do the work, and are they available now? Ask for names, seniority and start dates. Selling teams often pitch with senior people and staff with juniors. Put the named people in the contract and ask what happens if one leaves mid-project.
3. Where does my data go, and who can read it? You want a plain answer: which model provider, which region, whether your prompts train anyone's model, who on their side has production access. Vague reassurance here is the single biggest red flag in the whole list.
4. How will we know if the assistant is right? The answer should be an evaluation set built from your real questions with expected answers, scored before launch and re-scored after every change. "We'll test it as we go" means there is no measurement.
5. What happens when it does not know? Good assistants refuse, cite the document they used, and hand off to a person. A team that has not thought about refusal behaviour has not run one of these in front of real staff.
6. What is fixed and what is not? Ask for a fixed price and a fixed scope on the first phase, with the assumptions written down. Open-ended time and materials on a first AI build transfers all the risk to you.
7. Who owns the code, the prompts and the evaluation set? All three should be yours, in your repository, from day one. Prompts and evaluation sets are the real intellectual property in an assistant, and they are the easiest thing for a supplier to keep hold of.
8. What does it cost to run every month after launch? Model usage, hosting, monitoring and someone's time. A team that cannot give you a monthly running cost has not operated one of these at scale.
9. What would you tell me not to build? Any team worth hiring has a list. If everything you describe sounds like a great fit for their services, you are talking to a sales function.
10. What does the handover look like? Documentation, a runbook, a named owner on your side, and a support arrangement with response times. Ask to see a handover pack from a previous project with the client details removed.
Want a second opinion on a team you are already talking to? A 30-minute discovery call costs nothing and is not a pitch. Bring their proposal and we will tell you which parts we would push back on, where the scope is likely to move, and what we would build differently. You leave with a shortlist of questions to put to them. Book a discovery call.
How do you tell a builder from a talker?
The difference shows up in what they bring to the second meeting. Builders arrive with a narrowed scope, a rough evaluation plan and a question about your data access. Talkers arrive with a longer deck.

| Signal | Builder | Talker |
|---|---|---|
| Proof | A running system you can look at, with its limits described | Logos, awards, headcount |
| Team | Named engineers, CVs, availability dates | "Our expert team" |
| Scope | One process, fixed price, written assumptions | A phased transformation programme |
| Data | Asks about permissions and retention in meeting one | Raises security at contract stage |
| Measurement | Offers an evaluation set and a pass mark | Promises accuracy as an adjective |
| Honesty | Tells you what not to build | Everything is a fit |
| Best for | Companies that want one process working in production | Companies that need a board-level strategy artefact |
| Avoid if | You genuinely need a market-wide strategy review first | You need something live this quarter |
We built SofiaHR's candidate assessment platform this way. The useful detail for a buyer is not the technology. It is that the work started by testing ten different models against real assessment data before a line of product code was written, because the choice of model was a measurable question rather than a preference.
What should an AI assistant actually cost?
Two numbers to hold in your head. UK market daily rates for a senior AI engineer run from £600 to £1,200, and a middle AI engineer from £450 to £850 (SoftBlues rate card, UK market comparison). A first proof of concept for an assistant, two months, one to two engineers with part-time senior and project management support, prices at around £20,000.
That range is the honest one for a first phase with a defined scope. Anything materially cheaper is usually a demo with no evaluation layer and no integration work. Anything materially more expensive on a first phase is usually a discovery exercise wearing a build budget.
| Phase | Typical duration | Typical team | What you should get |
|---|---|---|---|
| Proof of concept | 2 months | 1 to 2 engineers, part-time senior and PM | Working prototype on your data, evaluation set, a written recommendation |
| Production build | 3 to 6 months | 2 to 4 engineers, QA, PM | Integrated assistant, monitoring, runbook, handover |
| Run and improve | Ongoing | Fractional engineer plus support | Monthly evaluation scores, model updates, a named owner |
For the fuller picture on budgets and rate structures, we set it out in AI consulting costs in the UK. If you are earlier than that and want to get the first phase right, how to run an AI proof of concept that reaches production covers the phase this post prices.
What are the red flags?
Six things that reliably predict trouble.
No named engineers in the proposal. You are buying capacity from a bench, and you will get whoever is free.
No evaluation plan. Without one, "it works" is a matter of opinion, and opinions change when the person who liked the demo leaves.
Accuracy claimed as a percentage with no test set attached. A number with no denominator is marketing.
A refusal to discuss what happens after go-live. Assistants drift as models update and documents change. A team with no run phase has not lived with one.
Discovery priced as a large fixed fee with no build commitment. Useful discovery ends in a costed build decision, not a report.
Everything is a fit. The strongest signal in a sales conversation is a supplier telling you which part of your idea to drop.
Which questions should your own side be ready to answer?
Vetting runs both ways, and the answers you cannot give are usually what delays the project.
If you cannot answer the first and the last of those, fix that before you brief anyone. It changes the price and it changes who you should hire.
Frequently asked questions
How long does it take to build an AI assistant? A useful first version against your own documents takes about two months. Production, with integrations, monitoring and a handover, usually adds three to six months depending on how many systems it touches.
Should I hire an in-house AI engineer instead? If you expect a steady pipeline of AI work for the next two years, yes, eventually. For a first assistant, an in-house hire means you learn the mistakes on your own payroll. Most companies get better value bringing in a team that has already made them, then hiring once the pattern is proven.
What size team do I need? One to two engineers for a proof of concept. Two to four plus QA and a project manager for a production build. If a supplier proposes more than that for a first assistant, ask what each person does in week one.
Do I need my own data scientists? No. Modern assistants are built on hosted models, so the work is retrieval, integration and evaluation rather than model training. You need engineers who can measure output quality, not researchers.
How do I check a supplier's claims? Ask to speak to a client whose system is still running, and ask that client what broke and how quickly it was fixed. Also ask the supplier for an anonymised handover pack. Both requests are ordinary, and reluctance to meet them tells you something.
What if the assistant does not work? Agree the exit before you start. A first phase should end in a written recommendation that includes the option to stop. We put a money-back guarantee on our 90-day production commitment for exactly this reason.
Can this run on Claude, or do I need something custom? Most business assistants are better built on a hosted model such as Claude than on anything custom. We run six of our own departments this way, which is documented in our own Claude operating system. Where a Microsoft-native shop is already standardised on Copilot, we will say so.
SoftBlues is a registered Anthropic Partner Network member, and a registered partner with Google Cloud and Microsoft. We build AI assistants for UK and Ireland companies with 50 or more knowledge workers, fixed price, in production in 90 days, money back if it fails. We use what we sell before we sell it, which is why the questions above are the ones we expect to be asked.
If you are choosing a team and want the shortlist pressure-tested, book a discovery call.
See it in production
Systems we have built and run for clients, with the numbers that came out of them.
Related Articles

Build vs Buy AI: What UK Mid-Market Companies Should Build in 2026 (and What They Should Not)

Shadow AI at Work: What UK Companies Should Do About Unapproved AI Tools
