Does the workflow deserve automation?
Start with a measured baseline: volume, cycle time, quality, cost, exception rate and business consequence. If the work has no stable outcome metric, there is no defensible ROI case.
Readiness before autonomy. Economics before theatre. A production-minded framework for leaders who need to decide where an agent belongs, how much authority it should have, what it must prove, and whether the economics close before a pilot becomes another expensive demo.
The model is only one component. Production readiness lives in the workflow around it: ownership, data, identity, permissions, evaluation, observability, recovery, human authority and business measurement.
A technically impressive agent can still be a bad investment or an unacceptable operational risk. Treat value, operability and control as separate gates; scale only when all three hold.
Start with a measured baseline: volume, cycle time, quality, cost, exception rate and business consequence. If the work has no stable outcome metric, there is no defensible ROI case.
The workflow needs usable context, callable systems, stable interfaces, explicit state transitions, retries, stop conditions and a way to recover when the world differs from the happy path.
Identity, least privilege, approvals, evals, logs, incident ownership and rollback are not paperwork around the agent. They are the mechanism that turns probabilistic behavior into accountable operations.
Score one real workflow, not “the company.” Choose the statement closest to current reality. The result is directional—not an audit—and stays in your browser.
Scoring: 0 = absent · 1 = informal · 2 = defined · 3 = evidenced and operational.
The useful question is not “How much will AI save the company?” It is “For this workflow, after automation quality, human review and run cost, what capacity or cost value remains?”
Scenario model only. “Capacity value” is time returned × loaded hourly cost; it is not automatically cash savings or revenue.
Do not assign one risk level to “the agent.” Reading a record, drafting a reply, changing a price and releasing a payment are different authorities with different failure costs.
The system retrieves, summarizes, drafts or recommends. A person performs the action.
The agent prepares the action with evidence; a human explicitly approves before execution.
The agent may act inside explicit policy, value, identity and rate limits; exceptions escalate.
The agent handles long-horizon work with durable state, controls and periodic human governance—not unrestricted independence.
A strong first deployment is measurable, bounded, frequent enough to learn from, technically reachable and recoverable when it fails. Customer or money-moving authority should come later unless controls are already mature.
| Workflow pattern | Value signal | Operational risk | Suggested first autonomy | Proof before scale |
|---|---|---|---|---|
| Internal research & synthesis | Expert hours, response time, decision throughput | Lower | Assist | Citation quality, coverage, hallucination rate, expert acceptance |
| Support triage & routing | Queue time, first-response time, escalation load | Lower | Propose + approve → bounded | Correct routing, missed-risk rate, handoff quality, customer outcome |
| Sales/account preparation | Prep time, meeting quality, rep capacity | Lower | Assist | Factual accuracy, CRM freshness, rep acceptance, conversion proxy |
| Contract pre-review | Legal cycle time, reviewer load, clause coverage | Medium | Propose + approve | Recall on material issues, false negatives, reviewer override reasons |
| IT / ops remediation | MTTR, ticket volume, downtime avoided | Medium | Bounded execute | Rollback success, change failure rate, permission boundaries, incident rate |
| Outbound customer communication | Cycle time, service volume, customer satisfaction | Medium | Propose + approve | Policy compliance, factuality, tone, complaint and escalation rate |
| Pricing, credit, payment or entitlement change | Processing cost, fraud/risk reduction, cycle time | High | Human/policy gate required | Legal basis, deterministic controls, value limits, dual control, full audit trail |
The goal of a pilot is not to prove that a model can generate plausible output. It is to produce enough operating evidence to decide whether the workflow deserves production authority and additional capital.
A practical enterprise control model can borrow from multiple frameworks. They answer different questions: organizational governance, lifecycle risk, agentic security and legal obligations.
Useful as a lifecycle risk structure. NIST emphasizes continuous risk management, measurement and evidence—not a one-time checklist.
Official NIST source ↗Organizational requirements for establishing, implementing, maintaining and continually improving an AI management system.
Official ISO source ↗Security guidance for agent goal hijack, tool misuse, identity and privilege abuse, supply-chain risk and other agent-specific failure modes.
Official OWASP source ↗As of September 2026, Article 50 transparency duties apply; high-risk obligations follow later under the current implementation timeline.
EU AI Act Service Desk ↗| Control question | Operational evidence | NIST lens | ISO 42001 lens | OWASP agentic lens |
|---|---|---|---|---|
| Who owns the system and its risk? | Named owner, RACI, stop authority, incident process | Govern | Leadership, roles, AIMS processes | Rogue-agent / cascading-failure containment |
| What may the agent do? | Tool inventory, service identity, permission scopes, value/rate limits | Map + Manage | Operational planning & controls | Tool misuse; identity & privilege abuse |
| What proves acceptable performance? | Representative evals, thresholds, adversarial tests, regression suite | Measure | Performance evaluation | Goal hijack and unsafe execution testing |
| Can you reconstruct an incident? | Context provenance, tool logs, approvals, outputs, timestamps, versions | Measure + Manage | Monitoring, measurement, documented information | Observability across agent/tool chain |
| Can the system fail safely? | Timeouts, checkpoints, idempotency, rollback, fallback and escalation | Manage | Corrective action & continual improvement | Cascading failures / unexpected execution controls |
This crosswalk is an engineering and management aid, not legal advice or certification guidance. Applicability depends on jurisdiction, sector, role and use case.
Usage is not value. A production scorecard needs leading signals that help you steer and lagging signals that prove whether the deployment improved the business.
The decision to stop, narrow or stay at assistive mode is a success when the evidence says autonomy would destroy value or accountability.
The business cannot agree what “good” means, exceptions dominate the workflow, or the process changes faster than it can be evaluated.
Required permissions are too broad, consequential actions cannot be gated, or failure cannot be reversed or meaningfully investigated.
Human checking, tool spend, latency, failure cost or change-management burden erase the modeled benefit. Use simpler automation—or keep the human workflow.
These are the questions that usually determine whether an agent initiative becomes a controlled operating capability or remains an impressive prototype.
It is the evidence that a specific workflow can give software bounded authority to observe, propose, act, verify and recover while remaining measurable, auditable and owned. It is a workflow property—not a generic company maturity score.
Start with the current workflow baseline. Subtract human review, rework, model/tool cost, infrastructure and material failure cost from the value of time, capacity, quality, risk reduction or revenue movement that the agent actually changes.
Only when the action is bounded, permissions are least-privilege, failure is detectable and recoverable, evals support the decision, monitoring is live, and the residual consequence fits the organization’s risk tolerance.
At minimum: service identity, scoped permissions, trusted context, typed tools, evals, policy gates, audit logs, cost/time limits, retries, rollback or fallback, incident ownership, and human authority for consequential exceptions.
High enough volume to learn, a measurable baseline, bounded steps, reachable systems, reversible consequences and clear escalation rules. Novelty is not a selection criterion; evidence velocity is.
From 2 August 2026, Article 50 transparency obligations apply and enforcement began for applicable rules. Under the current EU implementation timeline, Annex III high-risk obligations are scheduled for 2 December 2027 and regulated-product high-risk rules for 2 August 2028. Always verify role- and use-case-specific applicability.
This page is the decision layer. The existing library carries the architecture, enterprise sequencing, memory and systems-engineering detail behind it.
The page is vendor-neutral by design. These sources anchor the governance, security, economics and regulatory concepts; the decision framework and scoring model are ImageFirm synthesis.
Four functions—Govern, Map, Measure and Manage—with continuous risk management and measurement across the AI lifecycle. NIST notes that AI RMF 1.0 is being updated.
Open source ↗Cross-sector companion resource for incorporating trustworthiness considerations into design, development, use and evaluation of generative AI systems.
Open source ↗International standard for establishing, implementing, maintaining and continually improving an AI management system.
Open source ↗Peer-reviewed guidance on agent-specific security risks including goal hijack, tool misuse, identity/privilege abuse and supply-chain vulnerabilities.
Open source ↗Current implementation timeline. As of September 2026, Article 50 transparency requirements apply; the current schedule places Annex III high-risk rules later.
Open source ↗Separates leading and lagging indicators and recommends balanced measurement across adoption, value, governance and risk.
Open source ↗Architecture patterns and a strong emphasis on matching technical complexity to business value rather than defaulting to multi-agent designs.
Open source ↗Current 2026 enterprise perspective on centralized identity, policy enforcement, visibility and governance as agent deployments scale.
Open source ↗The strongest first conversation is concrete. If you are responsible for a workflow with measurable value, real data, real systems and real consequences, bring the baseline and the constraints. We can pressure-test whether it should become an agent, a simpler automation—or remain human-led.
Selective engagement. No automated lead scoring. Your message goes to a human.