Why Your AI Project Needs a Business Analyst

Karim Khalifa

CTO Consulting

Senior Business Analyst

Karim is a Senior Business Analyst, Project Manager, and Process Engineer specialising in complex enterprise transformation. He works across financial services, telecommunications, retail, and government, managing requirements, aligning stakeholders, strengthening governance, and delivering measurable outcomes in multi-vendor environments.

Benefits Realisation, ROI and the Business Analyst in Enterprise AI

Most AI projects don’t fail because the technology doesn’t work. They fail because nobody figured out how the technology would turn into money. That makes the business analyst the most important person on the program.

Most AI Projects Never Pay for Themselves

For two years, boards funded AI on faith. That period is over, and the numbers are not kind. MIT’s Project NANDA research found that around 95 per cent of enterprise generative AI pilots have no measurable effect on profit. S&P Global found that more than four in ten companies dropped most of their AI projects during 2025 because they could not show a path to a return. RAND puts the overall failure rate above 80 per cent, roughly double that of ordinary IT projects.

Very few of those are technology failures. The models work; the demos land. What goes wrong sits around the technology. Nobody agreed what success meant. They claimed benefits but never measured them against a baseline. ‘Time saved’ never became anything a CFO recognises. Pilots that shone in a controlled test fell apart against real data, real workflows and real people.

So the enterprise AI problem is mostly a benefits problem. Benefits realisation used to be the chapter at the back of the business case that nobody read. Now it separates the organisations making money from AI from everyone else. The business analyst is best placed to hold it together, from the first idea to the review two years after go-live.

Why AI Breaks the Old Benefits Model

Traditional systems either work or they don’t. A payments platform processes the transaction, or it doesn’t. You write the requirement, test against it and sign it off. Models are not like that. A model is right a certain percentage of the time, and that percentage varies across customer segments, slips as the world drifts away from the training data, and can fall sharply when an upstream system or supplier changes. So acceptance stops being a yes or no and becomes a threshold: how good is good enough, and which mistakes can the business live with?

Benefits stop being single numbers and become ranges that depend on two things: how well the model performs, and whether people use it. The second gets far less attention than it deserves. Almost none of the value comes from the model. It comes from people changing how they work around it. A triage model whose recommendations case officers override every time gives you all of the cost and none of the benefit. Adoption also follows a J-curve: productivity dips before it lifts, because people need time to learn how to supervise something that is only probably right.

Costs behave differently too. On a normal project, spending largely stops at implementation. Here it doesn’t. Inference and licensing rise with usage, so success itself becomes a cost. Models need monitoring, retraining and governance for as long as they run, and data pipelines, the biggest and dullest line in any honest estimate, need constant care. Treat AI as a capital project with one-off costs and permanent benefits, and the case will look excellent going in and embarrassing coming out.

Two more things sit outside what older frameworks handle well. Model performance drifts without announcing itself, so the benefits attached to it drain away just as quietly, unless someone has connected model monitoring to the benefit measures. And value can go backwards. A biased outcome, an invented answer that someone acts on, or a privacy breach can cost far more than the efficiency gains are worth. Disbenefits belong inside the benefits model, not in a risk register nobody opens.

A Business Case That Survives Contact with Reality

Start with honest costs. Most of the effort in moving from pilot to production isn’t model development. It is data engineering, integration, governance and measurement. Add change management, training, assurance and the running costs above. A case that prices the model and waves at the rest isn’t conservative. It’s fiction.

On the benefit side, put every claimed benefit in one of five groups:

  • Efficiency: time and cost

  • Effectiveness: quality, accuracy, cycle time, first-pass yield

  • Experience: for customers and for staff

  • Enablement: things the organisation could not do before

  • Risk and compliance: losses avoided, obligations met more reliably

Then attach five things to each: a named business owner, a measurement method, a baseline, a target and a date. A benefit without a baseline and an owner is not a benefit. It’s a hope.

Now apply the test that most AI cases quietly fail. Minutes saved per task is the most quoted AI benefit and the least often realised. Saved time only becomes money when something converts it: capacity moves to revenue work or clears a backlog, overtime or contractor spend falls, or the organisation absorbs growth without hiring. That conversion must be designed, and one named person must be accountable for it. Most of the 95 per cent measured usage and declared victory. Usage is not value.

Because performance and adoption are both uncertain, model the benefits as scenarios rather than promises: conservative, expected and stretch, each tied to a stated level of performance and adoption. Nominate leading indicators up front, such as active usage, override rates, first-pass yield and rework volumes. They tell you early what the financial measures will say later, while intervening is still cheap.

And start every pilot with written success criteria and written stop criteria. Most pilots can’t declare success because nobody said what success would look like. Nobody defined failure either, which is how organisations end up with zombie pilots quietly consuming budget for years.

What This Asks of the Business Analyst

On a conventional project, the BA’s job is well understood: find out what is needed, analyse it, write it down, check it. On an AI initiative, it stretches because the BA is usually the only person translating in three directions at once. Between the business and the data science team. Between both of those and privacy, legal and ethics. And between what the business case promised and what will eventually be measured.

The work starts before the project does. A striking number of failed AI initiatives were never AI problems. They were rules problems, automation problems or broken processes dressed up as AI because that is where the funding was. So the first useful contribution is a blunt opportunity assessment: does this genuinely need a system that learns from data under uncertainty, does the data exist at the required quality, is the value worth the whole-of-life cost, and would something simpler do the job? Killing a bad use case there is one of the highest-return acts on the program.

During delivery, three familiar pieces of the job change shape. Data requirements become a specification in their own right, covering sources, lineage, quality thresholds, refresh frequency, consent and licensing limits and retention.

Acceptance criteria become business decisions rather than technical settings. Tighten a fraud model, and you catch more fraud but block more genuine customers. Loosen it, and you do the reverse. Where to sit on that trade-off is a commercial and reputational decision, not a data science one. The BA turns the statistics into consequences the business can argue about properly, then records the decision and its reasoning.

And human-in-the-loop design becomes a deliverable. Which outputs are actioned automatically? Which go to a person, and at what confidence level? Who may override, and how are overrides captured and fed back? How is accountability preserved, so that ‘the model decided’ never becomes an answer anyone offers a regulator?

Running through all of it is benefits stewardship. The BA sets the baselines before the build starts, the step most commonly skipped in AI delivery and the one that makes every later claim impossible to prove. The BA also holds the thread from objective to benefit, from benefit to the business change that produces it, from that change to requirements and from there to the evidence that the benefit arrived. When scope pressure hits, that thread shows exactly which benefits a descoping decision gives away.

The role then continues past go-live, which is where AI departs most sharply from tradition. Someone has to connect model telemetry to benefit measures, run realisation reviews, spot the leak when drift sets in and recommend whether to retrain, expand, constrain or retire. Without that work, benefits that were real in month three can quietly disappear by month 18.

The Artefacts That Carry the Thread

Shaping produces an opportunity assessment, a business case built on a benefits dependency network that maps objectives to benefits, benefits to the changes that release them and those changes to the capability underneath, a data readiness assessment, privacy and responsible-AI impact assessments and a short experiment brief for each pilot with its hypothesis, thresholds and kill criteria.

Delivery adds process models that quantify the operational change the case depends on, a requirements package including the data specification, model acceptance criteria and an evaluation framework covering fairness, robustness and explainability as well as accuracy, human-in-the-loop procedures, non-functional requirements extended to auditability, latency, security and cost per inference and a traceability matrix binding benefits to requirements to evidence.

Sustainment needs a benefits realisation plan, tracking reports that put model performance and adoption next to the financial measures they drive, and review papers that give governance a real decision: scale, hold, remediate, or retire.

It is a longer list than a conventional project needs. Each item answers a question that would otherwise be answered by assumption, and assumptions are what failure statistics are made of.

Neither Agile nor Waterfall Quite Fits

Waterfall assumes you can specify the outcome in advance and deliver to that specification. With AI, it can’t. No responsible practitioner can promise, before the work is done, that a model will reach a given performance level on a given dataset. A fixed-price commitment to 95 per cent accuracy is a contractual fiction, and the procurement team that writes it is buying a dispute. Waterfall also puts all the value at the end, exactly when drift and adoption risk are highest.

Agile strains elsewhere. Data engineering and enterprise integration carry lead times that resist two-week increments. Regulated environments need privacy assessments, security accreditation, ethics review and records compliance that can’t wait for a later sprint. And sprint velocity is easy to mistake for progress towards value, which is the same activity-versus-outcome confusion the failure statistics punish.

What works is a deliberate hybrid. Run discovery and model development as timeboxed experiments with thresholds and kill criteria set in advance. Run delivery on an Agile cadence. Wrap both in a staged investment framework that releases funding against evidence at each gate. Done then means something partly statistical: the model meets its agreed operating point on held-out data, passes its fairness and robustness checks and has monitoring in place.

For the business analyst, that means hypotheses and evaluation criteria rather than exhaustive up-front specifications, benefits kept linked to the backlog so sprints don’t drift away from value and enough comfort with governance to write the gate papers.

The Skills That Matter Now

The capability this demands extends the classical BA toolkit in three directions.

The first is analytical. Not the ability to build models, but real fluency in what precision, recall, false-positive rates, drift and evaluation design mean, enough to translate them into business consequences and ask a data science team better questions. Plus enough statistical judgement to tell a real effect from noise in a benefits report.

The second is commercial. Benefits management is a discipline in its own right, and BAs should know it as well as they know elicitation techniques. Add enough financial modelling to defend a whole-of-life cost model in front of a CFO, and scepticism about vendor benchmarks achieved under demonstration conditions, about extrapolating from a pilot team of enthusiasts to a workforce of realists and about any benefit expressed in hours with no route to dollars.

The third is human and institutional. AI succeeds or fails on adoption and trust, which puts facilitation, stakeholder brokerage and change capability at the centre of the role. The BA is often the only person the operations manager, the data scientist, the privacy officer and the union delegate will all speak to candidly. Add responsible-AI fluency: privacy law, anti-discrimination obligations, sector rules, records and accountability requirements and the assurance frameworks now settling into place. In a growing number of sectors, the question is no longer whether an AI system performs, but whether the organisation can demonstrate that it performs fairly, lawfully and accountably. The person who writes that demonstration is usually the business analyst.

Value Is Designed, Not Discovered

The organisations in the successful minority did not buy better models. They managed value with more discipline. They qualified their use cases ruthlessly, baselined before they built, defined success and failure before they started, designed the mechanisms that convert capability into cash and outcomes and kept measuring long after the launch communications had faded. None of that is glamorous. All of it is benefits realisation.

AI has not diminished the business analyst’s role. It has raised the stakes. On a conventional project, weak benefits practice costs an organisation some credibility at the post-implementation review. On an AI program, it is the difference between compounding genuine returns and funding an expensive education.

The thread from a benefits dependency network drawn at the start to a realisation review held years after go-live is long, unfashionable and decisive. The business analyst holds it. Organisations investing in AI capability would do well to invest, with equal seriousness, in the people who can prove it was worth it.

Next
Next

AI Roadmaps for CTOs: From Hype to Real Business Value