ImageFirm The FirmStart an enquiry
Sheet A-15 · Agent engineering · Interactive

The Agent Go-Live Gate

Twenty-four questions your agent has to answer before it is allowed to act without a person checking each run. Six of them are hard stops. Score it honestly, get a verdict, and leave with the three things to fix first.

24 questions6 gates6 hard stops~8 minutes0 data leaves this page

Autonomy is a decision someone signs, not a setting that drifts on.

Most agents do not fail at the demo. They fail in month three, when the person who watched every run has moved on, the prompt has been edited four times, and nobody can say who approved the version that is now sending the emails. The gate makes the hand-over from supervised pilot to unattended operation an explicit act by a named owner — with the conditions written down before the decision, not reconstructed after the incident.

Bring the person who owns the outcome and the person who built the agent. Half the questions are about intent, half about mechanics, and the gaps live where those two have not talked.

  • ScoreEach question is No (0), Partly (1) or Yes (2). Maximum 48.
  • HardSix questions are hard stops. Any hard stop below a full Yes caps the verdict at not cleared, whatever the total.
  • Tiers0–23 not ready · 24–35 supervised pilot · 36–43 limited autonomy with sampled review · 44–48 cleared for unattended operation.
  • Honest"Partly" is the right answer more often than "Yes". A gate you pass by generosity is a gate you did not run.
  • PrivateAnswers stay in your browser. Nothing is sent, stored or tracked — the same rule we hold our agents to.

The six gates

G1

Purpose and authority

What the agent is for, and where a person keeps the final word.

0 / 8
01Hard stopThere is one written job statement for the agent: the outcome it must produce, how that outcome is measured, and an explicit list of what it must never do.
02Hard stopEvery action the agent can take is classified as read-only, reversible or irreversible — and irreversible actions cannot complete without a named person signing off.
03A single named person owns the agent's outcomes. Not a team, not a vendor — a person whose name is on the incident report.
04The people affected by the agent's decisions know an agent is involved, and know how to reach a human about it.
G2

Data and sovereignty

What the agent reads, where it writes, and who else can see it.

0 / 8
05You can list, today, every data source the agent reads and every destination it writes to.
06Hard stopPersonal or confidential data is minimised at the input, and the agent cannot send it to a model provider or third party you have not contracted with for that purpose.
07Prompts, retrieved documents and outputs are logged in a store you control, with a retention rule someone signed off on.
08You could replace the underlying model within one quarter without rebuilding the agent.
G3

Reliability and evaluation

How you know it works, and how you will know when it stops working.

0 / 8
09There is a held-out evaluation set built from real cases with expected outcomes, and the agent's pass rate on it is a number you can quote.
10Hard stopThe pass rate exceeds a threshold the outcome owner agreed to before the agent was built, not after seeing the result.
11The agent is tested against malformed and adversarial inputs, including instructions hidden in the content it retrieves.
12When the prompt, a tool, or the model version changes, the evaluation runs automatically and a regression blocks the release.
G4

Guardrails and failure handling

What happens when the run goes wrong, and how far it can go before it is stopped.

0 / 8
13Hard stopEvery run has a hard budget — steps, tokens, time and money — and the agent stops when any limit is hit.
14When the agent is uncertain it escalates to a person instead of guessing, and the threshold for 'uncertain' is defined and tested.
15Every tool call is validated on input and on output, and each tool carries the minimum permissions the job needs.
16Hard stopThere is a documented kill switch that any on-call person can operate in under a minute, and it has been exercised.
G5

Human oversight and accountability

How people stay in the loop after launch, not only before it.

0 / 8
17Humans review a sampled share of runs on a fixed cadence, and the sample size was chosen for statistical reason, not convenience.
18Human overrides and corrections are captured and fed back into the evaluation set.
19An incident process exists: who is told, within what time, what is written down, and who decides whether the agent keeps running.
20The agent's decision in any given run can be explained to the affected person in plain language, after the fact, from the logs.
G6

Operations and economics

Whether it can be run, paid for and switched off by people other than its builder.

0 / 8
21Cost per successful outcome is measured and compared with the process it replaced, including the cost of human review.
22Latency, error and escalation rates are on a dashboard that a named person looks at on a fixed cadence.
23A named person other than the original builder is trained to operate and change the agent.
24A decommissioning plan exists: what happens to the data, the integrations and the people who came to rely on the agent if it is switched off.

Questions people ask before running it

What is an agent go-live gate?

A go-live gate is a fixed set of conditions an AI agent must meet before it is allowed to act without a person checking each run. It sits between the pilot, where a human watches everything, and unattended operation, where the human only samples. The gate makes the hand-over an explicit decision by a named owner instead of something that drifts.

How is the score calculated?

Twenty-four questions, each answered No (0), Partly (1) or Yes (2), for a maximum of 48. Six questions are hard stops: if any of them is not a full Yes, the verdict is capped at 'not cleared for unattended operation' regardless of the total. A high score with a failed hard stop means the foundations are good and one specific thing is missing.

How long does it take, and what do I need?

About eight minutes if you know the agent. You need the person who owns the outcome and the person who built it in the same room, because half the questions are about intent and half are about mechanics. Nothing you enter leaves your browser.

When should the gate be re-run?

On every change to the prompt, tools, model version or data sources, and on a fixed cadence — quarterly is a sensible default — even when nothing changed, because the world around the agent does.

Is this a legal or compliance assessment?

No. It is an engineering and governance readiness check drawn from ImageFirm's own agent work. It maps naturally onto the human-oversight, logging and risk-management expectations that regulators are converging on, but it does not replace legal advice for your jurisdiction and sector.

Provenance

The twenty-four conditions are distilled from ImageFirm's own agent builds and rollouts — the same questions we put to ourselves before an agent of ours runs without a hand on it. They are deliberately vendor-neutral: nothing here depends on a particular model, framework or cloud. The hard stops are the six failures we have seen cause real damage; the rest cause expensive embarrassment. Revised when our practice changes, not on a content calendar.

0 of 24 answered
Work with ImageFirm

Bring us the scorecard with your agent's name on it.

If the verdict was not the one you wanted, the gaps are now specific and named. We close them the same way we opened them: a person in final authority, the method in the open, and a written record of where the machine stops.

Start an enquiry