The half of the loop nobody prepares for

An interview loop is symmetric in theory and wildly asymmetric in practice. The team runs a structured process with rubrics, calibration and multiple assessors. You get eight minutes at the end of each conversation and a vague instruction to ask questions. Almost everyone spends those minutes on team size, tech stack and what a typical day looks like — questions whose answers are the same at a team doing excellent work and at a team burning a budget on demos that will never ship.

This guide is the other half. Fifteen questions, grouped into five themes, each with the phrasing you would actually say out loud, what a strong answer sounds like, what a weak answer sounds like, and — the part that matters — what the weak answer predicts about your first six months. It is not an interrogation script. It is a diagnostic, and like any diagnostic it is only useful if you know how to read the result.

  • Five themes, three questions each. Production reality, evaluation practice, unit economics, failure culture, and the shape of your actual job.
  • Every question is answerable by someone who has shipped. Difficulty answering is the finding, not rudeness on your part.
  • Weak answers predict specific futures, and those predictions are more useful than the answers themselves.
  • Timing decides everything. Almost none of these belong in the recruiter screen.
  • A scorecard turns fifteen answers into one decision, with an honest adjustment for stage.
  • Standing is a prerequisite. You get straight answers in proportion to the credibility you bring.

Why a wrong choice costs more in an AI role than in general software

Every job change carries risk. Three things make the AI version of that risk unusually sharp, and they compound.

The field has an unusually high ratio of well-funded teams doing genuinely unproductive work. Capital arrived faster than competence, which is nobody's fault but is a fact you should price in. A team can hold a substantial budget, a serious brand, a room full of clever people, and eighteen months of activity that produced no system any user depends on. In ordinary software this is rare, because the artefacts are legible — a service either serves traffic or it does not. In AI work, the demo is unusually persuasive and unusually far from production. A notebook that answers ten curated questions well can look, to a non-technical stakeholder and sometimes to the team itself, like a product that is nearly finished. It is not nearly finished. The distance between that notebook and something a support organisation in Manchester or a lending operation in Hyderabad can rely on is where the entire job lives, and some teams never cross it.

Titles are non-standard, so the words on the offer letter tell you almost nothing. As of 2026, "AI Engineer" is attached to at least four different jobs: a software engineer wiring model APIs into an existing product, an applied scientist doing retrieval and fine-tuning work, a platform engineer running inference infrastructure, and — commonly — a prompt-and-demo role that reports into marketing or innovation and ships nothing durable. "Machine Learning Engineer", "Agent Engineer", "Forward Deployed Engineer" and "AI Solutions Architect" are similarly unstable. Two roles with the same title at two companies on the same road in Bengaluru can share almost no daily reality. The only way to know which one you are being offered is to ask questions specific enough that the answer cannot be generic.

And the cost of getting it wrong is measured in learning-curve time, not money. This is the part candidates underweight. A year at a team with no production system is not a neutral year. You come out of it with slide decks, a handful of prototypes, and no answer to the question every subsequent interviewer will ask, which is what you shipped and what happened when real users met it. Twelve to eighteen months at the steepest part of the curve is the single most expensive thing you can spend, and it does not appear anywhere in the compensation discussion. A modestly paid role at a team that ships is worth more than a well-paid role at a team that demos, and the gap widens every quarter. The system design round you are preparing for is testing whether you have that production experience; the reverse interview is how you find out whether the job will give it to you.

Watch out

Funding, brand and headcount are not evidence of production maturity, and in this field they are only loosely correlated with it. Some of the most capable AI work happening as of 2026 is inside unglamorous companies with real users and real constraints, while some of the least productive is inside teams whose logos are on every conference banner. Judge on the answers below, not the letterhead.

How to ask without sounding adversarial

These questions work when they read as enthusiasm about the work and fail when they read as an audit. The difference is almost entirely timing, framing and follow-up discipline, and all three are learnable.

Timing. Match the question to someone who can answer it. A recruiter or agency contact cannot tell you about regression gates, and pressing them makes you look as though you cannot read a room. Save the operational questions for the hiring manager and, above all, for the future-peer conversation — which is usually the most candid thirty minutes in the entire process, because the person on the other side will have to work with the consequences of your answer as much as you will.

Conversation What to ask here What to avoid here
Recruiter screen Which team, which product surface, is anything live, who the role reports to Anything requiring engineering detail — they cannot answer and both of you know it
Technical rounds One or two questions attached to something that came up in the problem you just discussed Broad process questions — the interviewer is scoring you and the clock is short
Hiring manager Production reality, the shape of the first project, decision rights, stakeholders, budget ownership Stacking all fifteen; pick four that matter most to you
Future peer Evals, incidents, on-call, data access, how model decisions actually get made Anything that asks them to criticise their manager by name
After the offer Written follow-up on anything that stayed vague — entirely normal in both markets Reopening a question that was already answered clearly; it reads as distrust

Framing. Attach each question to something the interviewer has already said, and phrase it as what you would need in order to be effective rather than as a standard they must meet. Some phrasings that work, verbatim:

  • "You mentioned the summarisation feature is live — what does the traffic look like on a normal day?"
  • "I have been burned before by shipping something I could not measure. How would I know whether a change I made improved things?"
  • "Where does the compute budget for this team sit, and who signs off on it? I ask because it tends to shape what is realistic."
  • "What is the last thing one of these systems did in production that you did not want it to do?"
  • "If I joined in six weeks, what is the first thing you would want me to work on, and what would make it hard?"

Reading the room. Watch three things while the answer comes. Speed — a practitioner reaches for a specific example within a sentence or two, while someone constructing an answer starts with a general principle and works towards a hypothetical. Specificity — numbers with units and dates, a named system, a named person, an actual incident. And hedging quality: genuine practitioners qualify their own numbers unprompted, because they know precisely where the measurement is soft. If an interviewer deflects, take it gracefully and move on; the deflection is your data and pressing for a second attempt costs more than it yields.

Pro tip

Ask the same question to two different people in the loop, separately, and compare. Not to catch anyone out — the mismatch itself is the finding. If the hiring manager describes a weekly eval review and your future peer has never attended one, you have learned something no single answer could have told you.

Theme one: does anything actually reach production?

Start here, because every other theme is conditional on it. A team with nothing live cannot have a meaningful eval practice, cost model or incident history, and questions about those will produce polite fiction.

Question 1 — What is live right now, and who uses it?

Ask it like this: "What have you got running in production at the moment, and who is on the other end of it?"

Strong answer. A named system, a named user population, and a rough sense of volume. "The claims triage assistant has been live since February for about ninety operations staff, roughly four thousand cases a week." The interviewer may not have the numbers exactly and will say so, which is fine and in fact reassuring.

Weak answer. A list of capabilities rather than systems, in the future or continuous tense. "We are building an agentic platform for enterprise workflows." Or an answer where every named user is internal to the team that built the thing.

What it predicts. A first six months spent building things that are evaluated by whether they impress in a review, not by whether anyone uses them. You will not accumulate production experience, which is the currency you are actually there to earn, and your next interview will be harder than this one.

Question 2 — Of everything the team has built in the last year, what fraction shipped?

Ask it like this: "Looking back over the last year — of the things the team started, how many ended up in front of real users?"

Strong answer. An honest ratio, usually unflattering, with reasons. "We probably started eight things and three are live. Two died because the latency was never going to work and one got overtaken by a vendor feature." Teams that ship are comfortable discussing what they killed, because killing things is part of shipping.

Weak answer. "Everything we build ships" — which either means nothing gets started without a guarantee, or that "ships" has been quietly redefined to include internal demos. Equally weak: an answer that cannot distinguish between a pilot, a proof of concept and a product.

What it predicts. If the ratio is genuinely near zero, expect a rhythm of enthusiastic starts and quiet abandonments. The specific damage is motivational rather than technical: you will stop investing properly in work you expect to be cancelled, and that habit is unpleasantly hard to unlearn.

Question 3 — What is the oldest AI system here that is still running?

Ask it like this: "What is the longest-running thing you have in production, and what has it taken to keep it alive?"

Strong answer. Something at least a year old, described with the specific weariness of maintenance. Model deprecations survived, prompt regressions after a provider update, a retrieval index that had to be rebuilt. This answer is gold, because maintenance stories cannot be fabricated by a team that has never done maintenance.

Weak answer. Everything is under six months old at a company that has been working on AI for two years. That combination means things are being replaced rather than maintained — usually because nothing reached the durability threshold.

What it predicts. You will learn to build but not to operate, and operating is the harder and rarer half. It also predicts a particular political dynamic where new initiatives get resourced and existing systems quietly rot, which becomes your problem the moment you inherit one.

Theme two: how do they know it works?

A team without an evaluation story has no feedback loop, which means quality is decided by whoever spoke last in the review meeting. You will be debugging by vibes, and you will lose arguments to people with more seniority and less information.

Question 4 — How would I know whether a change I made improved things?

Ask it like this: "Suppose I change a prompt or swap a model next month. What tells me whether that was an improvement?"

Strong answer. A concrete mechanism, named. An offline suite with a golden set, an online metric, or ideally both, plus an honest account of the gap between them. "We have about four hundred cases with human-labelled outcomes, and we watch the escalation rate weekly. The offline set catches the obvious regressions and misses the subtle ones."

Weak answer. "We try it and see how it feels." Or a benchmark score with no relationship to the product. Or — the most common one, and the most misleading — a model-as-judge score with no human-labelled anchor underneath it, which measures agreement with a grader nobody has validated.

What it predicts. Six months of unfalsifiable disagreements. Every quality question resolves by seniority rather than evidence, and you will spend real time re-litigating changes that were fine. It also predicts silent regressions, because nothing is watching. Our guide to building evals that agents cannot game is worth reading before you take a role where you might have to build this practice from nothing.

Question 5 — Is there a regression gate in CI, and has it ever blocked a release?

Ask it like this: "Do model or prompt changes go through any automated check before release? Has it ever actually stopped something?"

Strong answer. Yes, with an example of it firing and someone being annoyed about it. The annoyance is the proof. A gate that has never blocked anything is either newly built or set so loose it is decorative.

Weak answer. "It is on the roadmap." Or a gate that exists but which anyone can override without a record. Or a suite that runs but whose failures are habitually ignored because it is known to be flaky — which is worse than no suite, because it consumes credibility that the eventual real suite will need.

What it predicts. Release anxiety as a permanent condition. Every deploy becomes a judgement call made by whoever is most confident, and you will develop the defensive habit of changing as little as possible, which is precisely the habit that stops you learning.

Question 6 — Who owns the eval suite, and when was it last extended?

Ask it like this: "Who looks after the evaluation set? How often does it get new cases added?"

Strong answer. A named person or a rota, and a live process for turning production failures into new cases. The best version: "Anything that goes wrong in production becomes a test case before the fix ships." That single habit is the difference between a suite that improves and a suite that decays.

Weak answer. Silence, or a name followed by "though they have moved to another team". An unowned eval suite is a decaying asset that everyone still cites in meetings — which is more dangerous than no suite, because decisions are being made against a measurement nobody trusts enough to maintain.

What it predicts. Either you inherit the suite by default in month three, or quality drifts and nobody notices until a customer does. Both outcomes are survivable, but you should choose them knowingly rather than discover them.

Avoid

Accepting "we use an LLM as a judge" as a complete answer to any of theme two. It is a reasonable technique and a poor foundation on its own. The follow-up is: what is the judge validated against, and when did a human last check whether it agrees with people? If nobody has checked, the number is a comfort, not a measurement.

Theme three: who pays, and how much?

Cost questions feel commercial rather than technical, which is why candidates skip them. They are the best available predictor of whether the project will still exist next year, and you should know the answer before you sign.

Question 7 — What does an average request cost you?

Ask it like this: "Do you have a sense of what a single request or task costs to serve?"

Strong answer. A number with a unit, however rough, and awareness of the spread. "Around eleven rupees a case, but the long tail is four or five times that and it is the tail we worry about." A team in London might frame the same answer in pence per query. Either way, the answer exists and someone has looked at it.

Weak answer. "We have not looked at that." Or a total monthly spend with no per-unit breakdown, which means nobody can tell you whether the system gets cheaper or more expensive as it grows.

What it predicts. A cost review in the next two to four quarters that arrives as a surprise, followed by an emergency optimisation sprint under time pressure — the worst conditions for that work. Our guide to LLM unit economics and cost per task covers the model you will end up building in a hurry.

Question 8 — Is there a budget, or an open tap?

Ask it like this: "How is the model spend budgeted? Is there a number the team works within, and who owns it?"

Strong answer. A named owner and a rough envelope, with a sensible attitude to overruns. Even better: an answer that includes what happens when a team exceeds it, because that reveals whether the constraint is real.

Weak answer. "Nobody has said anything about cost." An open tap is not generosity, it is deferred scrutiny — and deferred scrutiny always arrives, usually attached to a headcount decision. Watch particularly for spend sitting inside a cloud commitment or provider credits that expire, because that is a cliff with a date on it.

What it predicts. A project whose economics have never been defended, meeting a finance function for the first time during a downturn. Teams that cannot answer this get cancelled in the next cost review, and the engineers find out at the same time as everyone else.

Question 9 — Has cost ever changed a technical decision here?

Ask it like this: "Has the cost of a system ever made you build it differently — route to a smaller model, cache more aggressively, cut a feature?"

Strong answer. A specific story with a trade-off in it. Routing simple cases to a smaller model, caching a retrieval step, dropping a reranking pass that was not paying for itself. The presence of a real trade-off proves the constraint is felt at engineering level rather than discussed at management level.

Weak answer. "We always use the best model available." That is not a quality standard, it is an absence of a cost model, and it usually coexists with the weak answers to questions 7 and 8.

What it predicts. A team that has never practised the engineering discipline of cost, being asked to acquire it overnight. That is a genuinely hard skill to build under duress, and the version of it done in a panic tends to damage quality in ways nobody measures — see theme two.

Every article here is written by a Verified Builder. Want your name on the next one?

AI Tech Connect lists AI engineers, founders and researchers across India and the UK — and the people hiring browse it to find them. Adding your profile is free.

Become a Verified Builder →

Theme four: what is the failure culture?

Every team that runs AI in production has had something go wrong. The only variable is whether they noticed, what they did, and whether they can talk about it. Vagueness in this theme is the loudest signal in the whole loop.

Question 10 — What is the last thing an AI system here did that you did not want it to do?

Ask it like this, and then stop talking: "What is the last thing one of these systems did in production that you did not want it to do?"

Strong answer. A specific incident, told without drama. A retrieval bug that surfaced one customer's document to another. An agent that retried a write and created duplicates. A tone failure in an outbound message that someone escalated. The teller will usually be slightly embarrassed and entirely factual.

Weak answer. "Nothing really springs to mind." At a team with genuine production traffic, that answer means nobody was watching, nobody told anyone, or nobody thinks it is the sort of thing one mentions. All three are the same finding.

What it predicts. The first serious incident on your watch will be handled without precedent, under pressure, in front of an audience. You will be improvising a response process at the worst possible moment.

Question 11 — Was there a postmortem, and what changed because of it?

Ask it like this: "What happened afterwards? Did anything change in how you build as a result?"

Strong answer. A written review, a small number of concrete changes, and at least one that made the system less capable in exchange for being safer. The willingness to accept a capability cost is the tell — it means safety is a real constraint rather than a stated value.

Weak answer. A fix with no review. Or a review where the conclusion was that someone should have been more careful. Individual blame as a root cause means the process will not improve, because nothing was learned that generalises.

What it predicts. Recurring incidents of the same shape, and a culture where you carry personal risk for systemic problems. That risk lands hardest on whoever joined most recently and understands the system least — which, for your first six months, is you.

Question 12 — What are the kill switches, and who can pull them?

Ask it like this: "If a system started behaving badly at three in the afternoon, what would you actually do? Who can stop it?"

Strong answer. A specific mechanism and a specific authority. A feature flag, a rate cap, a rollback path, and an on-call engineer who does not need permission to use it. Better still, an answer that includes blast-radius limits designed in advance — spend caps, write scopes, human approval above a threshold.

Weak answer. "We would roll back the deploy" for a system whose behaviour changed because a provider updated a model rather than because anyone deployed. Or a stop mechanism that requires waking a specific senior person, which means it will not be used at three in the afternoon on a Friday. Our guide to designing agents that fail safe covers what a good answer looks like from the inside.

What it predicts. Long incidents. Not more incidents necessarily, but slower ones, where the time between detection and containment is measured in hours because containment requires a conversation.

Theme five: what would you actually do here?

The first four themes describe the environment. This one describes your job, and it is where offers most often diverge from reality.

Question 13 — What is the first project, and why has it not been done yet?

Ask it like this: "If I started in six weeks, what would you want me working on? And what has stopped it happening so far?"

Strong answer. A specific piece of work with a specific obstacle, honestly named. "The retrieval quality on the policy corpus is poor and nobody has had two clear months to fix it." The second half is the valuable half — it tells you whether the blocker is capacity, which you can solve, or something structural, which you cannot.

Weak answer. "We will figure that out once you are here." Or a project that has been attempted twice already by people who have since left, described without curiosity about why.

What it predicts. Either drift for the first quarter while you look for something to own, or inheriting a problem whose real obstacle is political and was never going to be solved by hiring an engineer.

Question 14 — Who decides which model and which vendor, and how?

Ask it like this: "How do model and vendor decisions get made here? Is that an engineering call, a procurement call, or somewhere in between?"

Strong answer. A named process with engineering input, and awareness of the constraints. Data residency requirements for an Indian financial services client, or a UK public sector customer's procurement rules, are legitimate constraints and a team that names them is a team that has done the work. An answer describing a real evaluation — options considered, criteria, who decided — is what you want.

Weak answer. A vendor chosen by someone two levels up with a relationship, and no route to revisit it. Or the opposite failure: no decision process at all, with each engineer choosing independently and nobody tracking the resulting spend or exposure.

What it predicts. How much of your job is building versus persuading. If model choice is settled elsewhere and unrevisitable, expect to spend a meaningful share of your time working around constraints you had no part in setting — and to explain, repeatedly, why something is hard.

Question 15 — What is the on-call and data-access reality?

Ask it like this: "What does support look like for these systems out of hours? And on day one, what data would I actually be able to see?"

Strong answer. A clear rota with a clear scope, and an honest account of data access including its limits. "You will have access to the anonymised evaluation corpus from week one; production records need a separate approval that takes about a fortnight." Constraints described precisely are a good sign — it means someone has thought about them.

Weak answer. "It is pretty informal" for a system with external users, which usually means the last person to touch something owns it forever. On data: an inability to say what you would see, which frequently means the answer is very little, and that debugging will happen through intermediaries.

What it predicts. The texture of your daily life more than any other question here. No data access means you cannot investigate failures yourself, which caps how fast you can learn regardless of how good the team is. Informal on-call means unbounded interruption, which caps how much deep work you get. Both are worth knowing before you accept, and both are easier to negotiate as terms of the role than to fix afterwards.

Turning fifteen answers into one decision

Score each theme green, amber or red, then read the combination rather than the total. Some reds are survivable and instructive; some combinations are a warning you should take seriously.

Theme Green Amber Red
1. Production reality Named systems, named users, something over a year old Live but young; one real system and several pilots Nothing live, or "live" means internal demos
2. Evaluation Golden set, an owner, gates that have fired, prod failures become cases A suite exists but is thin, unowned or advisory only No mechanism; quality decided by opinion
3. Unit economics Cost per task known, budget owned, cost has changed a design Total spend tracked, no per-unit view, nobody alarmed No number, no owner, and no curiosity about either
4. Failure culture A specific incident, a blameless review, a change that cost capability Incidents handled ad hoc; reviews sometimes, informally No incident recalled, or individual blame as root cause
5. Your actual job Specific first project, honest obstacle, real data access, bounded on-call Project shape clear, access or on-call vague No defined work, no data access, unbounded support load

How to read the pattern. Theme one red is disqualifying on its own — if nothing reaches production, the other four themes are hypothetical and the role cannot give you what you came for. Theme four red is close behind, because it is a culture signal rather than a maturity signal, and culture does not improve because you joined. Three or more reds across five themes describes a team that has not yet met the consequences of its own work; you would be joining to absorb those consequences on someone else's timeline.

Amber is normal and often the best available. A team that is honest about a thin eval suite and wants help building one is offering you the most valuable thing on the market, which is a well-defined gap you are qualified to close. That is a much better proposition than a team claiming green everywhere.

The honest adjustment

Stage changes the meaning of a red. A four-month-old seed-stage company in Cambridge or Chennai with no golden set and no cost model is red on themes two and three for entirely legitimate reasons — nothing is in front of enough users for either to be measurable. What you are judging there is whether the team can describe the gaps accurately and say when they intend to close them. A forty-person, two-year-old, well-funded team with the same two reds is telling you something completely different: that it has had both the time and the money to build these practices and has not. Same colour, opposite conclusion.

What you need ready in return

The reverse interview has a prerequisite that nobody mentions: you get straight answers in proportion to the standing you bring into the room. Ask a hiring manager how they measure quality when you have not demonstrated that you could build such a measurement, and you will get a polite, general answer. Ask the same question having already shown work where you built an eval suite, watched a cost number, and handled something going wrong in production, and you get the real answer — because the person answering has recognised you as a peer who will find out anyway.

This is not about seniority or years. It is about legibility. The candidates who get candid answers are the ones whose work is visible before the conversation starts: a public record of what they have built, what it did, what broke and what they changed. Not a CV, which asserts, but a body of work, which demonstrates. A recruiter in Manchester or a founder in Bengaluru who can read your projects before the call arrives at that call already treating you as someone whose questions deserve real answers.

Concretely, three things buy you that standing. A public artefact for at least one system you took to production — what it does, who used it, what the failure modes were, what the cost profile looked like. An honest account of something that went wrong, which is disproportionately persuasive because so few people publish it and it signals exactly the failure culture you are screening for in theme four. And a durable place for all of it to live that is not a document you attach to applications — a profile someone can find, check and share internally before they have met you. Our guide to building an AI engineering portfolio as proof of work covers what to include and, more usefully, what to leave out.

The logic is direct. These fifteen questions ask a team to be specific about production, measurement, money and failure. A team is far more likely to be specific with someone who has already been specific in public. Verifiability runs both ways, and the candidate who has made their own record checkable is the one with the standing to ask that a team make theirs checkable too.

From a verified Builder

"I used to save my questions for the end and ask about team culture. Now I ask what broke last, and I ask it of the engineer, not the manager. Twice it has changed my mind about an offer I was ready to accept — and once it turned a company I was lukewarm about into the one I joined, because the answer was so precise and so unembarrassed."

— Arjun, Verified Builder · Bengaluru, India

The question behind the fifteen questions

Everything above collapses into one thing you are trying to learn: has this team met reality yet? Reality is users who behave unexpectedly, quality that drifts when a provider updates a model, a finance function that asks what this costs, and a Tuesday afternoon when something does the wrong thing at scale. A team that has met reality talks about it in specifics, without embarrassment, because it has already had these conversations internally. A team that has not talks in capabilities, roadmaps and potential.

You can learn from either, but only one will teach you the things that compound. The specifics are load-bearing: numbers with units, systems with names, incidents with dates. When you get those, you are almost certainly talking to a team that will make you better. When you get abstractions in reply to concrete questions, take the answer seriously — it is the most reliable information the loop will give you, and it is being offered freely.

And when you do accept, the questions do not stop being useful. The five themes are a good map of what to look at in your first ninety days, because the gaps you identified from outside are the ones you will be best positioned to close from inside — while you can still see them clearly, before they become normal.