The job in one paragraph
You are the entire AI function of an organisation that decided it needed one, and most people around you formed their view of the field from a conference talk or a vendor deck. Nobody will tell you what to do, because nobody knows. The pattern you set in month one is the one everyone expects in month twelve. Take requests and return impressive prototypes, and you spend two years as a demo factory. Measure something, ship something narrow and report honestly on both, and you own a defensible capability.
The plan below is deliberately unglamorous. Week one is access and a baseline. Week two is a use case that can be scored. Week three is one small ship, instrumented. Week four is making it legible to people who do not read code. Onboarding guidance circulating in 2026 for applied AI roles tends to describe the daily loop as prompt, eval, trace, fix, re-eval, and to recommend model access on day one, a locally runnable eval harness, and sandboxed access to production traces. Those are secondary sources rather than research — but they name the right three things to fight for in week one.
It helps to know why the seat exists. The employer-side finding is the one that matters here: NASSCOM reports that 50 per cent of Indian employers cite a skills mismatch, most acute in applied, production-ready skills such as RAG and MLOps — against a base in which around 16 per cent of Indian IT professionals are AI-skilled, a figure attributed to the Ministry of Electronics and Information Technology. Demand for the role is not the constraint; applied skill is. That is a strong position and an exposed one at once, in India and the UK alike, and it is why the month below is built around evidence rather than enthusiasm.
| Week | Goal | Artefact you produce | How you know it worked |
|---|---|---|---|
| 1 | Access, and a measured baseline of what exists | A one-page baseline: volume, current method, quality, cost per case | You can state today's error rate and per-case cost with a date attached |
| 2 | Pick the use case, and decline the others visibly | A scored triage table covering every request you have received | Your sponsor has agreed the pick and seen the rejections in writing |
| 3 | Ship something narrow and instrument it | An eval harness with a golden set, plus cost-per-task logging | A change you make either passes or fails the gate without a debate |
| 4 | Make it legible and set the next quarter | A one-page memo: baseline, ship, cost, eval, recommendation, refusals | A non-technical executive can repeat your conclusion accurately |
Week 1 — Access, and a baseline nobody has measured
The first week is not for building. It is for acquiring the conditions under which building is possible, and measuring the world as you found it.
What to demand on day one
Ask for all four in your first two days, in writing, with a reason against each. Requests made in week one read as onboarding; in week three they read as complaints. If you asked good questions before accepting — our guide to the reverse interview covers fifteen worth asking — you already know which will be contested.
| What to ask for | Why it matters | What to do if you do not get it |
|---|---|---|
| Model API access with a real budget — keys, a named spend limit, and permission to exceed it for a day | Without a budget you unconsciously optimise for not spending money, which is the wrong optimisation in month one | Ask for a small explicit number rather than an open tap. Refused entirely, that is a signal about the mandate — record it and escalate once, calmly |
| Sandboxed access to production traces — a redacted export of real inputs and outputs, not synthetic samples | Everything useful you learn about failure modes comes from inputs you would not have invented | Ask for a hundred redacted examples with a defined scope. If data protection is the blocker, get a named owner and a date, and hand-build a golden set meanwhile |
| A place to run evals — a machine or CI runner where a suite can execute repeatedly against real keys | An eval you can only run manually on your laptop will not survive a busy month | Start on your laptop with everything version-controlled, and put the CI runner in the week-four memo as a specific, costed ask |
| A named business owner for the first use case — one person accountable for the outcome, not a committee | Without an owner, success is unfalsifiable and every stakeholder gets a veto at the end | Do not start until you have one. This is the single item worth delaying work over, because building for an unowned outcome wastes the whole month |
Send the four asks as one short message, and copy the person who will be annoyed if they are not granted — not as a manoeuvre, but as a courtesy that creates a record. In week four you will want to say that traces were requested on the third and granted on the nineteenth, without sounding accusatory.
Measure what exists, even if what exists is a person
This is the most important act of week one, and it gets skipped because it feels like preparation, not work. Whatever you are asked to improve already happens somehow: tickets triaged by two people in Bengaluru who have done it for three years, product descriptions written by a copywriter in Leeds at a rate she could tell you if asked. That process is your baseline. Measure it now, or you will have nothing to compare against in week four, and your work gets judged on whether it impressed people rather than whether it helped.
The measurement need not be sophisticated. Sample eighty cases from the last month, record volume per week and what a case costs in staff time, then score them by hand against a definition of correct that you write down and your business owner agrees. That definition is worth more than any model you deploy this month, and it will be wrong in interesting ways — cheaper to discover now than in week eight. Send it as a single dated page before the week ends. People will correct you, and the number your work is measured against gets fixed while nobody has a reason to argue.
Week 2 — Find a use case that can actually be scored
By the start of week two you have a list of requests: a chatbot over the intranet, a tool that reads CVs, the thing a competitor showed at a trade show. That list is not a roadmap but a collection of enthusiasms, and your job is to turn it into a decision using criteria other people can see and argue with.
The three axes
Score every candidate on three things. Does it have a ground truth you can score? For a given input, can a competent person say whether the output was right? Classification, routing, extraction and retrieval score well; open-ended generation scores badly unless you can write a rubric, and "make it sound better" scores zero. Does someone own the outcome? A named individual whose objectives improve if this works — not a department, not a steering group. Does it fail safely? A misrouted ticket costs four minutes; a wrong figure in a quotation costs far more; a wrong answer about loan eligibility belongs in a different conversation.
| Candidate use case | Scoreable ground truth (0–3) | Named outcome owner (0–3) | Fails safely (0–3) | Total |
|---|---|---|---|---|
| Bengaluru fintech: support-ticket triage and routing — classify incoming queries and route to the right queue | 3 — twelve months of resolved tickets carry the correct queue as a label | 3 — the head of customer operations, whose handling-time target it moves | 3 — a misroute costs minutes, and a human reads every ticket regardless | 9 |
| UK retailer: product-copy generation pipeline — draft descriptions for a long tail of catalogue items | 2 — no single right answer, but a written style rubric plus paired human preference gives a usable score | 3 — the e-commerce merchandising lead, who owns catalogue coverage | 2 — bad copy is public, but every draft is reviewed before publication and is trivially revertible | 7 |
| Either market: "an assistant that can answer anything about our data" — requested by an executive sponsor after a conference | 0 — the question space is unbounded, so there is nothing to score against | 0 — everybody is enthusiastic and nobody's objectives change | 1 — plausible wrong answers about internal facts, with no review step | 1 |
Seven or more is a candidate; four or less is a demo, whatever its sponsor believes. For a middling score, propose the smallest work that would raise one axis — usually a fortnight of labelling that creates a ground truth.
How to say no, and how to say no upward
The biggest failure mode in this role is becoming a demo factory for whoever asked most recently, and it happens through twenty small accommodations rather than one bad decision. Saying no is therefore a routine habit, and it works best when impersonal. The mechanism is the scoring table: share the whole thing, including the requests you will decline, rather than issuing verdicts one at a time. A rejection inside a list of eleven scored candidates reads as process; the same rejection alone reads as a judgement of the asker.
Saying no upward needs a different shape. When it comes from a founder or the executive who sponsored your hire, do not reject it — reprice it. Say what it costs in weeks, say what stops, and ask them to choose: "that is roughly three weeks, and the ticket-routing work would not ship this quarter." It returns the decision to whoever has the authority to make it, and senior people are better at trade-offs than at being told an idea is bad.
The most expensive request arrives with "it should only take you a couple of days". Nothing involving a model, real data and a real user takes a couple of days, and accepting that framing is how a month disappears one afternoon at a time. Answer with your own written estimate, including the eval work.
Week 3 — Ship something small and instrument it
Week three is where the temptation is strongest: you now know the domain well enough to imagine something impressive, and the impressive thing is more fun to build. Build the narrow one. A boring, evaluable first ship establishes that AI work here is normal engineering with measurements attached. Narrow means one input type, one user group and a human in the loop: handle the easiest quarter of the queue, pass the rest through untouched, keep it behind a flag. The point is that a live system generates real traces of inputs behaving in ways you would never have invented.
The eval harness is your first real artefact
Before it goes near a user, write the harness. It need not be clever; it needs to run in one command and produce a number you can compare across changes. Start with the golden set from week one — forty to a hundred and twenty cases is plenty — and a scoring function so simple nobody can dispute it. Resist model-as-judge scoring until you have checked its agreement with a human on fifty cases.
# evals/run.py — the smallest thing that honestly counts as an eval harness.
import json, statistics
from pathlib import Path
from app.pipeline import answer # the function you are actually shipping
GOLDEN = Path("evals/golden.jsonl") # 40-120 hand-checked cases, version-controlled
THRESHOLD = 0.82 # today's pass rate; CI fails if a change drops below it
def score(case, output):
"""Return 1.0 or 0.0. Start with rules a colleague can verify by reading them.
Only add a model-as-judge once a human has agreed with it on 50 cases."""
if case["type"] == "routing":
return 1.0 if output["queue"] == case["expected_queue"] else 0.0
if case["type"] == "extraction":
expected = set(case["expected_fields"])
got = set(output.get("fields", {}))
return 1.0 if expected <= got else 0.0
return 0.0
def run():
results = []
for line in GOLDEN.read_text().splitlines():
case = json.loads(line)
output = answer(case["input"])
results.append({"id": case["id"], "score": score(case, output), "output": output})
pass_rate = statistics.mean(r["score"] for r in results)
failures = [r["id"] for r in results if r["score"] < 1.0]
print(f"pass_rate={pass_rate:.3f} n={len(results)} first_failures={failures[:10]}")
Path("evals/last_run.json").write_text(json.dumps(results, indent=2))
# The regression gate. This line is the whole point of the file.
if pass_rate < THRESHOLD:
raise SystemExit(f"REGRESSION: {pass_rate:.3f} < {THRESHOLD}")
if __name__ == "__main__":
run()
Commit the golden set beside the code, and make a new case part of the fix for every production failure. That one habit — nothing is fixed until the failure is a test case — separates a suite that improves for years from one nobody trusts.
Error analysis, from traces rather than from intuition
Once anything is live, read the traces. Not an aggregate dashboard — the actual inputs and outputs, in batches of fifty. Group failures into named categories and count them: retrieval misses, formatting, and cases where the human label was wrong and the system right. Those categories become your work queue, ordered by frequency rather than interest. Two sources repay attention in both markets: transliterated input breaks pipelines tested only in English, and an Indian support queue carries Hinglish, Tamil and code-switched text that eighty hand-picked cases under-represent; format variation — UK postcodes, GST identifiers, date orders — breaks retrieval quietly, in ways that look like model stupidity but are parsing.
Judging quality by the outputs you happen to see. Spectacular failures and satisfying successes are both rare; the mundane middle — slightly wrong, plausibly formatted, never escalated — is where the real error rate lives, and only systematic sampling finds it.
Cost per task, from the first day it is live
Instrument spend before anyone asks. The number that matters is not the monthly bill but the cost per completed task, with retries and failures in the numerator. Log it per request, tag it by feature. Building the model properly — margin, the long tail, ten times the volume — is covered in our guide to LLM unit economics and cost per task. The logging must exist from the start.
# obs/cost.py — record cost per task, not cost per month.
import json, time
from dataclasses import dataclass, asdict
from collections import defaultdict
# Price per 1M tokens in your billing currency. Keep this in config, not in code.
# NOTE: the numbers below are ILLUSTRATIVE PLACEHOLDER RATES as of September 2026.
# They are not any vendor's published pricing. Replace them with your own contracted
# rates before you report a cost figure to anyone.
PRICES = {
"small": {"in": 0.25, "out": 1.25},
"large": {"in": 3.00, "out": 15.00},
}
@dataclass
class Call:
task_id: str
feature: str # "support-triage", "product-copy" — the unit finance cares about
model: str
tokens_in: int
tokens_out: int
latency_ms: int
ok: bool # did the TASK complete, not did the API return 200
def cost_of(call):
p = PRICES[call.model]
return (call.tokens_in * p["in"] + call.tokens_out * p["out"]) / 1_000_000
def record(sink, call):
row = asdict(call) | {"cost": round(cost_of(call), 6), "ts": time.time()}
sink.write(json.dumps(row) + "\n")
def rollup(rows):
"""Cost per COMPLETED task, by feature. Retries and failures stay in the numerator."""
spend, done = defaultdict(float), defaultdict(int)
for r in rows:
spend[r["feature"]] += r["cost"]
done[r["feature"]] += 1 if r["ok"] else 0
return {f: round(spend[f] / max(done[f], 1), 4) for f in spend}
Week 4 — Make it legible, and set the next quarter
Everything you have done in three weeks is invisible to the people who decide whether this role continues. Week four is translation work, and a solo AI hire cannot delegate it.
The one-page memo
Write one page — not a deck, not a document with an appendix. A page can be read in a meeting nobody prepared for, forwarded intact and quoted accurately, which is what determines whether your work survives contact with the organisation.
ONE-PAGE MEMO — [feature name] — [date]
Owner: [you] Business owner: [named person] Status: [shipped / paused]
1. BASELINE (what it was before, measured in week one)
Volume ................ [n per week]
Current method ........ [e.g. two agents triaging by hand]
Quality baseline ...... [x% correct on 80 sampled cases, measured DD Mon]
Cost baseline ......... [staff minutes per case, or cost per case]
2. WHAT SHIPPED
Scope ................. [one sentence]
Explicitly NOT ........ [one sentence — the more specific, the better]
Live since ............ [date] Traffic: [n per week] Rollout: [% or which team]
Human in the loop ..... [what a person still checks, and when]
3. WHAT IT COSTS
Cost per completed task [amount, including retries and failures]
At current volume ..... [amount per month]
At 10x volume ......... [amount per month]
Breaks if ............. [the threshold that changes the economics]
4. WHAT THE EVAL SAYS
Golden set ............ [n cases] Pass rate: [x%] Gate: [threshold] in CI
Known failure modes ... [two or three, named and counted]
NOT covered by eval ... [be explicit — this line buys you trust later]
5. RECOMMENDATION FOR THE NEXT QUARTER
Do .................... [one thing, expected effect, how it will be measured]
Decide ................ [one decision I need from you, by when]
Not doing ............. [two or three declined requests, each with a reason]
6. WHAT I NEED
[budget / access / a named reviewer / annotation hours — with numbers]
Three details separate a memo that lands from one that gets filed. The baseline goes at the top: a result without a comparison is a number, not a finding. Section five names what you refuse as explicitly as what you propose — refusals written down become decisions, refusals kept in your head become resentments. And section four states what the eval does not cover: volunteering the limits of your own measurement tells a sceptical reader the rest of your numbers are honest.
Talking about non-determinism with people who have only shipped deterministic software
Your colleagues have spent their careers with systems that either work or have a bug. That a system is right eighty-four times in a hundred, and that a provider updating a model can change your quality without anyone deploying, is genuinely foreign; treating it as obvious is the fastest way to look evasive. Explain it once, early, in terms they already accept: a spam filter has a false-positive rate, a fraud model has a threshold trading one error against another. Then give the number, the failure modes and the review step that catches the worst cases. People accept probabilistic systems when the rate is plain and the containment is visible.
Be equally plain about the timeline. Vendor and consultancy material circulating in 2026 claims that teams building AI capability systematically report productivity gains of 25 to 55 per cent within ninety days. Those figures are not independently verified and the measurement is rarely disclosed, so do not let them become your yardstick — but assume your executives have read them. Month one buys a method, not a transformation.
Every article here is written by a Verified Builder. Want your name on the next one?
AI Tech Connect lists AI engineers, founders and researchers across India and the UK — and the people hiring browse it to find them. Adding your profile is free.
Become a Verified Builder →The traps
Five failure modes account for most of the unhappiness in a solo AI role. All are structural rather than technical, and all are cheaper to handle in month one.
The demo treadmill. You build something impressive, it is well received, someone asks for another, and the applause is pleasant. Six months later nothing is in production and nothing is measured. The tell is a calendar driven by upcoming presentations rather than a queue you control; the exit is the week-two scoring table, shared, plus a standing answer that trades any new demonstration against named work.
Inheriting a vendor decision made before you arrived. Reopening a platform or provider choice in month one is almost always a mistake: you spend all your credibility on a fight you cannot win with three weeks of evidence. Measure within the constraint, note in the memo where it costs you, and build the file. A quarter of documented cost or quality impact is an argument that a founder's preference is not.
Being made the AI policy owner without authority. This arrives disguised as a promotion: you decide what the organisation may do with these systems, without legal beside you and without power to stop anyone. Draft the document willingly — you are the only person who can — then insist on a named, recorded sign-off in legal, compliance or risk. Frame that as speed, not self-protection: a policy without an accountable owner cannot approve anything, and blocks your own work first.
Getting no annotation budget. Everyone agrees evals matter and nobody wants to fund the labelling that makes them possible, which quietly kills a first AI function: thin measurements lose arguments to confident opinions. Ask in small, specific units — "four hours a week of one domain expert for six weeks" is something a manager can grant, while "an annotation budget" goes to a committee. Do the first fifty yourself.
The single-person on-call problem. If you are the only person who understands the system, you are the only one who can respond when it misbehaves — unbounded by default. Design blast-radius limits into the first ship — spend caps, rate limits, a feature flag anybody on the existing rota can pull without your permission — then write the runbook while the system is still simple enough to describe in a page. The goal is that the first response never requires you specifically to be awake.
Why your external record matters more here
Here is the structural fact nobody says out loud at the offer stage. As the only AI person in the organisation, you have no internal peer review. Nobody checks your architecture, catches your blind spots, or can vouch for your judgement in a way that carries weight outside the building. Your manager can say you were good; they cannot say your retrieval design was sound, and the next employer knows it.
That makes your external evidence trail matter more here, because a sole AI hire's career risk is concentrated. If the function is restructured, if the sponsor leaves, if the budget moves, no team of colleagues can attest to what you built. Indeed's Hiring Lab put UK postings overall down 11 per cent since the start of 2026, but that is a market-level series and says nothing about the security of any individual role. The plainer point stands without it: restructuring is a normal thing that happens to new functions, and a new function of one is the easiest of all to unwind. You have whatever is legible from outside.
The hedge is discoverability, built from artefacts you already produce. The eval harness from week three, generalised and stripped of anything confidential, demonstrates exactly the skill NASSCOM's employers say they cannot find: applied, production-ready judgement rather than framework familiarity. That is why our guide to turning evals into portfolio proof of work treats them as the strongest single artefact a solo builder can show. Write it up as you learn it, not during a crisis.
The common failure here is assuming the work speaks for itself internally. It does, right up until the person who understood it changes jobs and everything you built exists only in a private repository and one departed manager's memory. Write up the method once a month — never the employer's data, only the approach. It costs an hour, and it is the part of the job that stays yours.
There is a practical reason to make that record findable, not merely to keep it: employers report low applicant volume while capable engineers find nobody sees their work. A Verified Builder profile on AI Tech Connect narrows that gap. No directory guarantees an outcome, but when your professional risk sits inside one organisation, a checkable record elsewhere costs little to maintain.
What thirty days can and cannot buy you
By the end of the month you should have five things: a dated baseline, a scored triage table showing what you chose and declined, one narrow system in front of real work, an eval harness with a regression gate, and a cost per completed task. That is not a transformed organisation, and anyone who promised one has made month two harder. It proves you can turn an unbounded ambition into something measurable and defensible.
You will also have a clear read on whether the role can give you what you came for. If access arrived, the sponsor engaged and the narrow ship reached real users, month two writes itself. If three weeks went on chasing permissions and every conversation was about a demonstration, that is information arriving early and cheaply. Either way, read the rubric you were hired against from the other side: our guide to hiring a team's first AI engineer sets out what your manager was told to look for.