What the failure data actually tells you
Start with the uncomfortable part. As of September 2026, the best-known numbers on enterprise AI outcomes are not close calls. MIT's NANDA initiative, in its report on the state of enterprise generative AI published in late 2025, found that 95% of generative AI pilots delivered no measurable impact on profit and loss. The RAND Corporation has reported that more than 80% of AI projects fail — roughly twice the failure rate of conventional IT projects. And a 2025 study from MIT Sloan found that 61% of enterprise AI projects were approved on projected return on investment that was never measured after launch.
Read those three together and a specific conclusion falls out, and it is not the one most engineers reach. The obvious reading is that enterprises are bad at AI. The useful reading is that approval is not the bottleneck — measurement is. A majority of these projects cleared the funding gate. They were approved, staffed and built. What they did not have was a number afterwards that anyone could point at. And a project with no measured result is not neutral in the next budget round; it is the easiest thing in the portfolio to cut, because nobody has to argue against evidence to cut it. They only have to shrug.
That reframes the entire task. Your job is not to win an argument in a room. Your job is to construct a project that can survive the twelve months after that room. The two are related but they are not the same, and optimising for the first at the expense of the second is how a talented engineer ends up with three cancelled projects and a reputation for enthusiasm.
One more figure, and it is the operational hinge of this whole article. One 2026 analysis found that projects with quantified success metrics defined upfront achieved a 54% success rate, against 12% for those without. That is a single-source finding and should be treated as directional rather than precise, but the direction is consistent with everything else here, and it points at something you personally control. You cannot fix your organisation's data foundations this quarter. You can absolutely write down the metric, the baseline and the reporting date before you start.
The recurring root causes cited across these studies are worth memorising, because your one-page case should visibly pre-empt each one: unclear definitions of success, weak data foundations, poor integration into real workflows, chasing technology rather than business outcomes, and fading executive sponsorship. Notice that exactly none of them is a modelling problem. Every one is an organisational problem that shows up as a technical failure eighteen months later.
The most dangerous phrase in an internal AI case is strategic importance. It gets you approved and it gets you cancelled, because a project justified on strategic importance has no number to defend and therefore competes purely on the sponsor's attention. Sponsors change roles. Numbers do not.
Who actually signs
Most engineers write their case for one person — the manager who will say yes — and are then ambushed by three others who were never in the room. The decision set for internal spend is almost always larger than the org chart suggests, and each member is optimising for something different. Map them before you write anything.
| Who | What they are actually optimising for | What makes them refuse | What to hand them |
|---|---|---|---|
| Budget holder (function or cost-centre owner) |
Hitting their own committed plan without surprises. Their bonus is usually attached to a number you can find out. | An open-ended cost, or anything that could go wrong loudly in front of their own boss. | A capped build cost, a stated run cost, and a written kill criterion so the downside is bounded. |
| Finance business partner | Defending the number when it is challenged upward. They personally own the forecast line your project would sit in. | Savings that cannot be banked, benefits with no owner, and any figure they cannot reproduce. | Your baseline method, their loaded rates rather than yours, and an explicit realisation factor. |
| Security, data protection and compliance | Not being the reason for an incident. They are a gate, not a stakeholder, and gates do not negotiate on outcomes. | Being told about the project after the design is fixed. Personal or regulated data with no stated boundary. | A one-page data-flow note in week one, before you need anything from them. |
| Platform or architecture owner | Not inheriting an unsupportable system when you move team. Consistency with what they already run. | A novel stack nobody else operates, with no observability and no named owner after handover. | A run-book sketch and a named maintenance owner in the case itself. |
| Executive sponsor | A visible win they can talk about in their own forum, at low personal risk. Their attention decays fast. | Nothing to report for two quarters. Sponsors disengage from silence, not from bad news. | A dated reporting rhythm with a first number inside ninety days. |
| The team whose work changes | Not being made worse off, and not being the subject of an efficiency story told about them. | Learning about it from a slide. Being measured by a metric they had no part in choosing. | Co-authorship of the baseline measurement. Involve them first, not last. |
Two structural points about this table. First, the last row is the one engineers skip and the one that most often kills delivery. A workflow team that was not consulted will comply with your pilot and quietly route around it, and your adoption number will be inexplicably poor. Poor integration into real workflows is one of the named root causes in the studies above; in practice it usually means the people doing the work were treated as a data source rather than as co-designers.
Second, the executive sponsor is the least reliable component of the whole system. Fading executive sponsorship is on the root-cause list for a reason: sponsors are promoted, reorganised and redirected, and a project whose only defence is a sponsor's enthusiasm has a half-life measured in quarters. The counter-measure is to convert sponsorship into structure while the enthusiasm is high — a budget line, a named finance partner, a signed measurement contract and a reporting date. Those survive a reorganisation. Goodwill does not.
Find out what your budget holder is personally measured on this financial year and write the first line of your case in that language. If they are measured on cost per transaction, lead with cost per transaction. If they are measured on cycle time or on backlog, lead with that instead. The same project can be framed three ways; only one of them is in the language of the person signing.
The one page they will actually read
Assume your case gets six minutes of genuine attention, some of it in a corridor. That is not cynicism, it is arithmetic about how many things a director is asked to approve in a quarter. Everything you need must fit on one page, in business language, with the numbers visible without scrolling. A longer document is not more persuasive; it is more skippable, and it signals that you have not decided what matters.
The structure below has seven blocks and no more. Note what is deliberately absent: architecture, model names, vendor comparisons and anything with the word framework in it. Those belong in an appendix that most readers will never open, and that is fine — the appendix exists so that the technical reviewer can find it, not so that the budget holder can.
ONE-PAGE CASE — [project name] [date]
Owner: [you] Budget holder: [name] Finance partner: [name]
1. PROBLEM, IN BUSINESS TERMS
Supplier invoice queries are resolved by hand. The team of
12.4 FTE handles 12,000 queries a month at a measured
average of 9.0 minutes each. Backlog at month end is the
reason 4% of invoices are paid late.
[No technology mentioned. If it cannot be said without
naming a model, it is not a business problem yet.]
2. COST OF THE STATUS QUO (measured, not estimated)
Baseline measured 14-25 Aug 2026, sample of 180 queries,
timed with the handling team. Method note attached.
Volume 12,000 / month
Handling time 9.0 min average, 14.0 min at p90
Loaded rate [finance's number, not yours]
Rework 6.8% reopened within 14 days
Annual cost of handling: [figure]. Owner of this figure:
[finance partner name] — they supplied the rate.
3. PROPOSED CHANGE (one sentence)
Draft the resolution and the reply automatically, and have
a handler review and send, for the query types where the
answer is already determined by data we hold.
4. COST TO BUILD
20 engineer-weeks, capped. No new headcount requested.
Slice: 2 query types out of 9, one region, one system.
5. COST TO RUN (annualised, the year after go-live)
Model spend [figure] + platform [figure] +
0.25 FTE maintenance + [N] hours a week of review sampling.
Total: [figure]. This line does not disappear later.
6. MEASUREMENT CONTRACT
Metric blended handling minutes per query
Baseline 9.0 min (see section 2)
Target under 5.5 min on the piloted query types
Measured by [ops manager], not the build team
Method the same timing script used for the baseline
Reported on D+30, D+90, D+180, to [sponsor] and finance
7. KILL CRITERIA (agreed now, in writing)
Stop at D+90 if any of the following is true:
- blended handling time is above 7.0 min
- handler acceptance of drafts is below 50%
- any incorrect payment reaches a supplier
On a stop, we retain the baseline measurement and the
evaluation set, and we write up why. Cost of stopping:
[figure], known in advance.
Section seven is the one that surprises people, and it is the one that most reliably gets a yes. Volunteering the kill criteria looks like weakness and is in fact the strongest move available to you, because it converts an unbounded risk into a bounded one. A budget holder is not afraid of a project that fails; they are afraid of a project that fails slowly and invisibly and cannot be stopped without a political fight. Hand them the off switch and you have removed the main reason to say no.
Section five is the one engineers most often omit, and omitting it is how a project gets approved and then resented. The build is a one-off; the run is forever. A case that prices twenty engineer-weeks and says nothing about the annual cost of keeping the thing alive will be re-litigated the moment finance notices a recurring line nobody planned for. If you want that conversation to go well, read our guide to forecasting your LLM bill before launch and put the resulting figure in the case yourself, with its assumptions written down.
The numbers you must own
There is one rule underneath this entire section: you cannot claim an improvement without a pre-measured baseline. Not a remembered baseline, not a baseline reconstructed from a dashboard after the fact, and definitely not a baseline supplied by the team that wants the project approved. If the baseline is measured after the system exists, the comparison is worthless and everyone in the room knows it, whether or not they say so.
The baseline has four components and they are all obtainable in a fortnight without any new tooling. Cost per task before the change: human minutes multiplied by the loaded rate your finance partner gives you. Volume: how many of these happen, per month, with the seasonal shape if there is one. Error and rework rate: what proportion come back, and what a failure costs when it escapes. Distribution, not just the mean: the p90 matters more than the average, because the long tail is usually where both the pain and the automation opportunity live. If your organisation logs none of this, sample it by hand — one hundred to two hundred real cases, timed with the people who do the work, with the method written down so somebody else could repeat it. Hand-counting two hundred cases is not beneath you; it is the single most fundable fortnight of work available in most organisations, because it is the only artefact in the whole project that nobody can argue with.
Then you need the unit economics of the thing you propose to build, and this is where engineers are strong and usually still get it wrong by being too optimistic. Tokens per task multiplied by price is the starting point, not the answer. You must add retries and timeouts, the evaluation calls that will run in production, embedding refreshes, and the oversized inputs that pull the mean well above the median. Our deeper treatment of LLM unit economics and cost per task works through the margin implications properly; what follows is the minimum viable version for a business case.
A worked example, end to end
Take a finance shared-services team handling supplier invoice queries — a workload that exists in almost identical form in a Bengaluru or Hyderabad global capability centre and in a Manchester or Leeds enterprise back office. Every rate below is a placeholder. Replace each one with the number your finance partner owns; using their rate rather than an invented one removes the first and cheapest objection anyone can make to your case.
| Line | UK enterprise example | India GCC example |
|---|---|---|
| Loaded handler cost (placeholder rate) | £36.00 per hour | ₹1,200 per hour |
| Handling cost before the change: 9.0 minutes | £5.40 per query | ₹180.00 per query |
| Handling cost after: 4.8 blended minutes (70% reviewed in 3 min, 30% unchanged at 9 min) |
£2.88 per query | ₹96.00 per query |
| Gross time saving | £2.52 per query | ₹84.00 per query |
| Model spend, including retries and eval sampling | £0.16 per query | ₹17.87 per query |
| Model spend as a share of the gross saving | 6.4% | 21.3% |
| Saving finance will bank (45% realisation) | £1.13 per query | ₹37.80 per query |
| Net benefit per query | £0.97 | ₹19.93 |
Two things in that table deserve more attention than the headline figures. The first is the realisation factor. Your gross saving is 2.52 currency units per query; the number finance will actually put in a plan is a fraction of that, because saved minutes are not saved money until something changes — a role is not backfilled, growth is absorbed without hiring, or a contractor line is reduced. Apply a discount yourself, in your own case, before finance applies a harsher one for you. Forty to fifty per cent is a defensible opening position and volunteering it buys you enormous credibility, because it demonstrates that you understand the difference between a benefit and a saving.
The second is the ratio of model spend to benefit, and it is the single most important dual-market observation in this article. In the UK example the model spend is 6.4% of the gross saving — a rounding error, and cost engineering barely affects the case. In the Indian GCC example, with a lower loaded human rate against identical token costs, the same architecture consumes 21.3% of the benefit. The arithmetic is the same; the conclusion is not. In a GCC, a chatty agentic design that would be invisible in a UK cost base can eat a fifth of the case before you have paid for a single engineer-week. If you are building in a lower-cost base, cost-aware design is not an optimisation you do later — it is a precondition of the case existing at all, which is why cost-aware evaluation belongs in the pilot rather than in a follow-up ticket.
The model, as code
Put the case in a file rather than a spreadsheet cell, so that when someone challenges an assumption in the room you can change one constant and answer them immediately. The following runs as-is and reproduces every figure quoted above.
# Internal AI business case — illustrative model.
# EVERY rate below is a PLACEHOLDER. Replace with the numbers
# your finance partner owns. Currency is one consistent local unit.
# --- Baseline: measure this BEFORE you build anything ---
VOLUME_PER_MONTH = 12_000 # tasks handled by the team
MINUTES_PER_TASK = 9.0 # measured from 180 timed cases
LOADED_COST_PER_HR = 36.00 # from finance, not from a job board
# --- The proposed system ---
AUTO_SHARE = 0.70 # share the system drafts confidently
REVIEW_MINUTES = 3.0 # human review time on drafted cases
FALLBACK_MINUTES = 9.0 # unchanged for everything else
# --- LLM unit economics (measure from real runs, do not guess) ---
IN_TOKENS = 45_000
OUT_TOKENS = 3_000
PRICE_IN_PER_M = 3.00 # USD per 1M input tokens
PRICE_OUT_PER_M = 15.00 # USD per 1M output tokens
RETRY_FACTOR = 1.12 # retries, timeouts, re-runs
EVAL_SAMPLE_RATE = 0.05 # share re-scored by an automated judge
EVAL_IN_TOKENS = 8_000
EVAL_OUT_TOKENS = 400
FX_TO_LOCAL = 0.79 # USD to local currency. Re-check this.
# --- Fixed costs and the honesty dial ---
BUILD_COST = 80_000 # 20 engineer-weeks at an internal rate
RUN_COST_PER_YEAR = 66_700 # platform + 0.25 FTE + review sampling
# (model spend is counted per task below)
REALISATION = 0.45 # share of saved minutes finance will bank
def cost_per_task():
t_in = IN_TOKENS * RETRY_FACTOR + EVAL_SAMPLE_RATE * EVAL_IN_TOKENS
t_out = OUT_TOKENS * RETRY_FACTOR + EVAL_SAMPLE_RATE * EVAL_OUT_TOKENS
usd = t_in * PRICE_IN_PER_M / 1e6 + t_out * PRICE_OUT_PER_M / 1e6
return usd * FX_TO_LOCAL
def net_benefit_per_task():
per_min = LOADED_COST_PER_HR / 60.0
before = MINUTES_PER_TASK * per_min
after_min = AUTO_SHARE * REVIEW_MINUTES + (1 - AUTO_SHARE) * FALLBACK_MINUTES
gross = before - after_min * per_min
return gross * REALISATION - cost_per_task()
def payback_months():
monthly = VOLUME_PER_MONTH * net_benefit_per_task()
surplus = monthly - RUN_COST_PER_YEAR / 12.0
return float('inf') if surplus <= 0 else BUILD_COST / surplus
def break_even_volume(horizon_months=12):
"""Monthly volume needed to repay the build over `horizon_months`."""
fixed = RUN_COST_PER_YEAR + BUILD_COST * (12.0 / horizon_months)
return fixed / (net_benefit_per_task() * 12.0)
print(f"cost per task {cost_per_task():.4f}") # 0.1604
print(f"net per task {net_benefit_per_task():.4f}") # 0.9736
print(f"payback (months) {payback_months():.1f}") # 13.1
print(f"break-even volume {break_even_volume():,.0f}") # 12,557
Trace the base case by hand, because you will be asked to. Effective input tokens are 45,000 × 1.12 plus 5% of 8,000, which is 50,800. Effective output is 3,000 × 1.12 plus 5% of 400, which is 3,380. At the stated prices that is $0.2031 per task, or £0.1604 at the stated conversion. Handling cost falls from £5.40 to £2.88, a gross saving of £2.52; at 45% realisation that is £1.134, and net of model spend, £0.97 per task. Across 12,000 tasks a month the project throws off about £11,683, against a fixed run cost of £5,558 a month, leaving £6,124 of monthly surplus. The £80,000 build therefore repays in about thirteen months, and the break-even volume for a twelve-month payback is 12,557 tasks a month — slightly above the 12,000 you actually have.
That is a deliberately awkward result and you should present it as one. The case does not pay back inside the financial year at current volume. It pays back in month thirteen and yields roughly £73,000 a year thereafter. Saying that plainly is far more persuasive than rounding the assumptions until year one looks positive, because every experienced finance partner has seen the rounded version and discounts it on sight.
Sensitivity: what actually moves the answer
Run the model with one assumption changed at a time and bring the table into the room. This is the part that converts you from an enthusiast into someone worth funding, because it shows which risks matter and which do not.
| Scenario | Net benefit per task | Payback on the build | What it means |
|---|---|---|---|
| Base case as written | £0.97 | 13 months | Positive, but not inside one financial year. |
| Volume triples to 36,000 a month | £0.97 | 3 months | Volume is the strongest lever. Pick the high-volume workflow. |
| Model price halves | £1.05 | 11 months | Barely moves it. Do not build the case on a price cut. |
| Automation share is 55%, not 70% | £0.73 | 25 months | A 15-point miss on acceptance nearly doubles payback. |
| Finance banks 25% of saved minutes, not 45% | £0.47 | No payback worth discussing | Settle the realisation factor before you build, not after. |
The lesson in those five rows is stark. Halving the model price — the thing engineers spend the most time optimising and the thing vendors talk about most — moves payback by under two months. A fifteen-point shortfall in how often handlers accept the draft nearly doubles it. And the realisation factor, a number set by a finance conversation you may never have had, can extinguish the case entirely. Spend your preparation time in that order: realisation, then adoption, then volume, then unit cost.
Bring the sensitivity table, not the base case. Presenting one confident number invites the room to attack the assumption behind it. Presenting five scenarios with the assumption named in each moves the conversation from is this number right to which of these scenarios do we believe — a discussion you can win, and one in which the budget holder becomes a participant rather than a judge.
The measurement contract
This is the highest-leverage idea in the article and the direct answer to that MIT Sloan finding about 61% of projects being approved on a projection nobody ever checked. The measurement contract is a short written agreement, settled before build starts, that fixes five things: the metric, its baseline value, the method by which it will be measured, the person who will run that measurement, and the dates on which it will be reported. It is not a dashboard, it is not a KPI in a strategy deck, and it is not a paragraph in a project charter. It is a page with signatures on it.
Three properties make it work. First, the metric is a business metric, not a model metric. Nobody outside your team can act on an F1 score or a retrieval hit rate; they can act on handling minutes, resolution rate, cost per transaction or days to close. Keep the model metrics — they are how you debug — but they are instrumentation, not the contract. Second, the measurer is not the builder. If your team both builds the system and reports its benefit, the number will be discounted by every reader, and correctly so. Hand the measurement to the operations owner or to finance, and make your job to supply the method and defend it. Third, the reporting dates are calendar dates agreed in advance, which is what converts a decaying sponsor's attention into a standing obligation that survives their departure.
MEASUREMENT CONTRACT — [project name] v1.0 [date]
Metric Blended handling minutes per query, on the
two piloted query types only
Baseline 9.0 minutes. Measured 14-25 Aug 2026 from a
sample of 180 timed cases. Method: [ref]
Target Under 5.5 minutes at D+90
Guardrails Rework rate does not rise above 6.8%
Zero incorrect payments released to suppliers
Cost per query stays under [ceiling]
Measured by [ops manager] — NOT the build team
Method The same timing script and sampling frame
used for the baseline. Re-runnable by either
party, producing the same number.
Reported on D+30, D+90, D+180, in writing, to the budget
holder, the finance partner and the sponsor
Decision rule At D+90: on target, we extend to the
remaining query types; between target and
the kill threshold, we hold and re-measure at
D+180; past the kill threshold, we stop.
Realisation Finance will recognise 45% of the measured
time saving in the plan. Agreed with [name].
Signed [budget holder] [finance partner] [build lead]
Getting this signed is a harder conversation than getting the project approved, and that asymmetry is precisely why it is valuable. Approval costs a sponsor nothing; a measurement contract costs them the option of quietly declaring success later. If your organisation will not sign one, you have learnt something important very cheaply — that the project is being funded for reasons that are not the stated reasons, and that its survival will depend on politics rather than results. That is worth knowing in week one rather than in month fourteen.
On the mechanics of the measurement itself: define the method so precisely that a sceptic could re-run it. Which cases are in the sample, over what window, excluded how, timed with which clock. Where the output is a judgement rather than a duration — quality of a drafted reply, for instance — you need a rubric and more than one reviewer, and our guide to building an evaluation suite with golden sets and judges is the practical reference for constructing something defensible. And once the system is live, attribute its cost properly rather than letting it disappear into a platform bill: per-feature cost attribution and showback is what lets you report a real cost per task at D+90 instead of an estimate.
"We got two projects approved in the same quarter. One had a signed measurement contract and one had an enthusiastic director. Eleven months later the director had moved to another business unit and that project was gone with no record it ever existed. The other one had a number at D+90, a number at D+180, and it is now funded as a permanent line. Same team, same technology, entirely different outcome."
— Arun, Verified Builder · Bengaluru, IndiaThe pilot that earns the budget
The purpose of an internal pilot is not to prove the technology works. You already know it works, and so does everyone else in the room; nobody in September 2026 needs convincing that a language model can draft a reply. The purpose is to move one business number on one real workflow, under conditions realistic enough that the result transfers. That is a different design brief, and it leads to a different pilot.
Scope it so that a negative result is still a cheap, useful result. That principle drives four concrete choices. Time-box it to something short enough that stopping is not a career event — six to eight weeks is usually the right shape for an internal slice, against the two weeks that suits an external fixed-price engagement. Pick a narrow slice with a willing owner: two query types out of nine, one region, one system, and an operations manager who actively wants the outcome rather than one who has been volunteered. Write the kill criteria before you start, in the case document, so that stopping is executing the plan rather than admitting defeat. And keep the measurement artefacts whatever happens — the baseline, the sampling frame, the evaluation set. Those outlive the pilot and are frequently worth more than the code.
Worth being explicit about the difference from external work, because the two are often conflated. Our companion guide on scoping a two-week paid pilot as a solo builder is about a contract with a client: the constraints are the fee, the fixed price and the exit. An internal pilot has no fee and no exit. Its constraints are your credibility, your team's time and the sponsor's attention span, and it is judged not on whether it was delivered but on whether the number moved. For the journey from a working internal pilot to something that survives production, our reporting on enterprise agentic AI going from pilot to production covers where that transition typically breaks.
Never pilot the hardest workflow first, however tempting the political logic of tackling the thing everyone complains about. The hardest workflow has the worst data, the most exceptions, the most opinionated stakeholders and the highest chance of an ambiguous result — and an ambiguous first result is the one outcome from which a project rarely recovers. Earn the right to the hard workflow with a clean, measured win on an easier one.
Headcount, contractors, or your own time
There are three ways to resource an internal project and they carry wildly different approval costs. Understanding that ordering is the difference between a case that clears in three weeks and one that dies in a planning cycle.
Your own time, reallocated. The cheapest thing to approve, because the money has already been spent. It is also the most commonly abused: an engineer volunteers evenings, delivers something, and establishes that the work costs nothing — after which asking for real resource becomes harder, not easier. If you use your own time, make it visible and finite. Ask your manager for a named allocation, say two days a week for six weeks, and have it recorded. Invisible effort produces invisible projects.
Contractors or an existing supplier. Operating expenditure rather than a permanent commitment, which makes it far easier to approve during a headcount freeze, and reversible, which makes the budget holder comfortable. The trade-off is that knowledge leaves when the contract ends, so the case must name who inherits the system and fund their handover time explicitly.
Permanent headcount. The slowest and most contested of the three, because it is a permanent cost approved on the basis of a temporary claim. This is why asking for headcount before proving the metric usually fails — you are asking someone to accept an ongoing liability against an unmeasured benefit, which is exactly the pattern the failure statistics at the top of this article describe. The sequence that works is invariable: measure the baseline, prove the delta on a narrow slice, then convert the proven delta into a role in the next planning cycle. A headcount request that opens with a measured result from a shipped pilot is a fundamentally different document from one that opens with a projection.
Both markets have their own version of this constraint. In a UK enterprise the binding limitation is often a headcount freeze rather than a cost limit, which is why contractor and supplier routes clear faster and why the benefit is best framed as absorbed growth, reduced backlog or shorter cycle time rather than as a role removed. In an Indian GCC the framing is more often cost per FTE and the centre's ability to take on higher-value work from the parent, so a case showing a lower unit cost per transaction — or a team moving up the value chain — carries further than a headcount saving, which may not be a welcome message in a centre measured partly on scale. Neither framing is more honest; they are answers to different questions, and using the wrong one is a straightforward unforced error.
One career note that engineers systematically underrate. The person who reliably converts a technical idea into funded, measured, delivered work is doing staff-level work regardless of their title, and it is usually the missing evidence in a promotion case built entirely on shipped systems. If that is the direction you are heading, our guide on moving from senior to staff AI engineer covers how this kind of organisational work is assessed and, more importantly, how to make it legible to people who did not see you do it.
Where internal cases go wrong
Six failure patterns account for most of the internal AI projects that never get funded, or get funded and then quietly disappear. Each maps to one of the root causes the research keeps identifying.
- The demo trap. A demo proves feasibility. It never proves value, and the gap between the two is the entire subject of this article. A convincing demo often makes funding harder, because it invites the response that the problem is already solved and does not need a budget. If you demo at all, demo the measurement, not the output: show the baseline, the sample and the method by which the difference will be established.
- Pricing the build and forgetting the run. Twenty engineer-weeks is a number people can approve. The annual cost of model spend, platform, monitoring, evaluation, drift management and a quarter of an engineer is a number they must live with, and discovering it after go-live poisons the relationship with finance for every project that follows.
- Letting a sponsor's enthusiasm substitute for a budget line. Enthusiasm is not funding. Until there is a cost centre, an amount and a named owner, you have permission rather than a project. Convert the first into the second while the enthusiasm lasts, because fading executive sponsorship is one of the most-cited causes of failure and the fade is rarely announced.
- Choosing a metric you cannot instrument. Customer satisfaction, decision quality and employee experience are real and largely unmeasurable inside a ninety-day window at a cost you can afford. Pick something you can count, on a sample you can define, with a method you can repeat. A modest metric that moves visibly beats an important metric that stays ambiguous.
- Piloting the hardest workflow first. Covered above, and worth repeating because the political pull towards it is strong. The workflow everyone complains about is the one with the worst data and the most exceptions.
- Treating security and compliance as a late-stage sign-off. A gate consulted in week one shapes a design; a gate consulted in week ten blocks one. Send a one-page data-flow note early, before you want anything, and the relationship is entirely different when you do.
Behind all six sits the pattern the studies keep naming: chasing technology rather than business outcomes. It is an easy trap for good engineers precisely because the technology is genuinely interesting and the business outcome usually is not. The discipline is to write the business sentence first, every time, and only then decide what to build.
Every article here is written by a Verified Builder. Want your name on the next one?
Getting an internal project funded, measured and paid back is the most portable work you will ever do, and it is completely invisible from the outside. A Verified Builder profile is where you make it visible — the method, not the employer. AI Tech Connect lists AI engineers, founders and researchers across India and the UK, and adding your profile is free.
Become a Verified Builder →Your first 90 days
A sequence you can start on a Monday. It is deliberately front-loaded with conversations and measurement rather than with building, which is the opposite of most engineers' instinct and the whole point.
Days 1 to 10 — find the workflow and the owner. List the candidate workflows in your area and score them on four things: volume, how measurable they are, whether the data you need already exists, and whether there is an owner who actively wants the outcome. Pick the highest-volume workflow that scores well on all four. Do not pick on interest.
Days 5 to 20 — measure the baseline. Sample one hundred to two hundred real cases. Time them with the people who do the work, alongside them rather than about them. Count errors and rework. Write the method down in one page so someone else could repeat it. Get your finance partner's loaded rate rather than estimating one. This is the fundable artefact even if nothing else happens, and it is what makes every later claim credible.
Days 15 to 25 — build the cost model and the sensitivity table. Use the code above. Measure tokens from real runs on a handful of representative cases rather than estimating from a prompt you wrote. Produce five scenarios, not one number. If you have never forecast a run cost before, do that work now rather than in the meeting.
Days 20 to 30 — write the one page and pre-brief. Never let the budget holder read your case for the first time in the meeting. Walk it round individually: the finance partner first, so the numbers are theirs by the time anyone challenges them; then security with the data-flow note; then the platform owner; then the budget holder. By the time it reaches a forum, it should already have three people in the room who have seen it and had their objection addressed. Meetings ratify decisions that were made beforehand.
Days 30 to 40 — sign the measurement contract. Before any build work starts. If it will not get signed, stop and find out why; the answer determines whether this project is worth your next quarter.
Days 40 to 85 — build the narrow slice. Two query types, one region, one system. Instrument cost per task from the first day rather than the last week. Hold a short standing checkpoint with the workflow owner so that adoption problems surface while they are still fixable.
Day 90 — report the number, whatever it is. On the agreed date, in writing, to the agreed people, against the agreed baseline. If it hit the target, ask for the next slice with the measured result attached and the sensitivity table updated with real figures. If it missed, apply the kill criteria you wrote, say so plainly, and write up what would have to be true for the answer to change. A project that stopped on evidence at day ninety is a considerably better line on your record than one that drifted for two years and was cancelled by a reorganisation.
That last point is the argument of this whole guide compressed into a sentence. Given that more than 80% of AI projects fail, by RAND's reckoning, the realistic goal is not to be certain you will succeed. It is to be the person whose projects always produce a defensible answer — because in an organisation where almost nothing is measured, being reliably measurable is the scarcest and most fundable quality you can have.