Essay
AI Generates; Systems Verify; Humans Decide
A risk-tiered operating doctrine for using AI at scale: let models create options, make systems test what can be tested, and reserve consequential judgment for accountable people.
“AI generates; systems verify; humans decide” is useful only if it changes the operating design.
It is not a demand to place a person behind every model response. That would turn oversight into a queue and make low-risk automation uneconomic. Nor is it permission to automate every decision once a confidence score crosses an arbitrary threshold. The doctrine assigns each kind of work to the mechanism best suited to it:
- AI generates candidates, classifications, scenarios, explanations, and proposed actions.
- Systems verify identities, authorities, arithmetic, completeness, limits, evidence, and policy conditions that can be tested consistently.
- Humans decide material trade-offs, exceptions, disputed facts, and consequential or difficult-to-reverse actions—and remain accountable for the decision environment.
The boundaries should move with risk. The doctrine is a control architecture, not a universal workflow template.
Start with consequence, not capability
Reported fact. The NIST AI RMF says risk-management intensity should reflect organizational risk tolerance. It calls for documented human-oversight processes, examination of the costs of AI errors, testing before deployment and during operation, and independent review where useful. NIST AI RMF Core
Reported fact. GAO’s accountability framework separates governance, data, performance, and monitoring. Its human-supervision guidance says the appropriate level depends on the system’s purpose and potential consequences; its monitoring practices include plans, acceptable drift ranges, traceable corrective actions, and conditions for scaling. GAO-21-519SP
These frameworks point away from a single control checklist. A meeting-summary draft and a capital-allocation recommendation may use similar underlying technology but do not create similar exposure.
Inference. Risk tier should be determined by the decision context, not the model’s brand or a generic label such as “agent.” Useful factors include:
- magnitude and distribution of potential harm;
- reversibility and time available to correct an error;
- whether an output stays internal or reaches a customer, market, regulator, or counterparty;
- sensitivity and authority of the input data;
- scale and correlation—one error versus the same error across thousands of cases;
- observability and time to detection; and
- whether the workflow can produce legal, financial, access, or payment effects.
Four tiers, not one ritual
An organization can translate those factors into a small number of operating tiers. The examples below are a design pattern, not a regulatory classification.
| Tier | Typical use | Verification | Decision and oversight |
|---|---|---|---|
| 0 — exploratory | Brainstorming, private drafting, non-production analysis | Basic data-loss and access controls | User judges usefulness; no operational side effect |
| 1 — bounded operations | Reversible classification or routing inside fixed limits | Schema, permissions, allowed actions, reconciliations, confidence and exception rules | System may execute in bounds; humans monitor samples and exceptions |
| 2 — material recommendation | Financial analysis, customer communication, compliance or risk recommendation | Authoritative-source checks, business rules, validation suite, complete evidence record | Named accountable person approves before external or material effect |
| 3 — high-consequence action | Irreversible payment, rights, safety, capital, or legal action | Independent validation, strict permissions, dual control, fail-closed gates, recovery testing | Restricted AI role; explicit authorization and tested override or shutdown path |
No tier removes governance. It changes the form and intensity. Tier 0 still needs appropriate data handling. Tier 1 needs monitored boundaries. Tier 2 needs decision evidence. Tier 3 may exclude generative output from the execution path entirely.
Open question. Which characteristic moves a use case up a tier: transaction value, affected population, data sensitivity, irreversibility, or a combination? If the escalation rule is not explicit, business pressure will quietly redefine “low risk.”
AI generates: widen the option set
Generation is valuable where the task benefits from search across possibilities: drafting a variance explanation, proposing a forecast narrative, mapping an exception to likely causes, or producing alternative operating plans. The model should create a candidate, not manufacture authority.
The intended-use contract should state what the model may produce, which sources it may use, what it must not do, and what constitutes failure. This keeps a fluent response from being mistaken for a supported conclusion.
Reported fact. NIST’s Generative AI Profile defines confabulation as confidently presented erroneous or false content. It recommends empirically evaluating capability claims, sharing pre-deployment test results with release authorities, reviewing sources and citations in outputs, monitoring overrides, and documenting risk decisions. NIST AI 600-1
Inference. The generator should usually be treated as an untrusted proposer. That is not a statement that every output is wrong. It is an architectural stance: eloquence, confidence, and internal consistency do not establish that a payment is authorized, a total reconciles, or a source supports a claim.
Systems verify: turn policy into executable tests
Verification belongs in systems when the condition can be expressed and observed reliably. Examples include:
- identity, role, and segregation-of-duties checks;
- schema, range, and referential-integrity validation;
- reconciliations to ledgers or other authoritative records;
- source existence, citation resolution, and version checks;
- transaction, exposure, and rate limits;
- duplicate and idempotency controls;
- allowed-tool and allowed-action policies;
- required-field and evidence-completeness gates; and
- logs, alerts, rollback markers, and exception routing.
A second model can add useful challenge, but model-against-model agreement is not automatically assurance. Shared training patterns, prompts, retrieval sources, or missing data can create correlated confidence. Whenever possible, pair probabilistic generation with a different verification mechanism: deterministic arithmetic, a policy engine, an authoritative database, a signed approval, or a sampled independent test.
Inference. “System verified” should mean that named checks ran against named evidence and produced recorded results. A generic score without calibration, threshold rationale, and failure disposition is an observation, not a control.
Humans decide: preserve judgment and accountability
Human review works only when the reviewer has information, time, competence, authority, and a real ability to disagree. A person clicking “approve” on a high-volume queue is not necessarily exercising judgment. The operating design should show the candidate, source evidence, failed and passed checks, uncertainty, prior overrides, and the consequence of each available action.
Reported fact, with an important scope limit. On April 17, 2026, the Federal Reserve issued SR 26-2, transmitting revised Federal Reserve, OCC, and FDIC model-risk guidance emphasizing a tailored, risk-based approach, effective challenge, validation, monitoring, clear roles, and documentation. The attachment explicitly excludes generative and agentic AI from its scope. It adds that a banking organization’s governance should guide the controls for tools outside the document’s coverage. Its relevance here is therefore an analogy for control architecture—not a claim that SR 26-2 directly governs a generative-AI workflow.
Reported fact. The IIA describes its 2024 AI Auditing Framework as principles-based guidance covering AI governance, management, controls, and internal audit, emphasizing reasonable assurance, transparency, traceability, and accountability.
Inference. Decision ownership and assurance ownership should not collapse into the same role. The person rewarded for throughput should not be the sole judge of whether controls are effective. Operations owns the decision; risk and control functions challenge the design; internal audit, where applicable, evaluates the governance and control system rather than operating it.
A capacity model for proportional control
Illustrative estimate (hypothetical). Assume a workflow processes 50,000 low-value internal classifications each month. Reviewing every item for 90 seconds would require 1,250 hours. A risk-tiered design might automatically process cases that pass bounded system checks, route 2%—1,000 cases—to five-minute human exception review, and independently sample 500 passed cases for two minutes each. That is about 100 review hours before investigation and governance overhead.
This is not a promised 92% saving. The assumed exception rate, review times, sampling plan, control efficacy, and cost of escaped errors are invented. The example makes a narrower point: identical human review is not the only path to oversight. Automation can preserve evidence on every item while people focus on exceptions and statistically or risk-selected samples.
If monitoring finds a rising override rate, new failure mode, data shift, or concentrated loss, the workflow should move up a tier, narrow its authority, or stop. Scaling is a decision to be earned by evidence, not a default reward for high volume.
The minimum evidence packet
For a consequential output, retain enough evidence to reconstruct what happened:
- intended use and risk tier;
- input sources, relevant versions, and data timestamp;
- model, prompt or policy version, and tools invoked;
- generated candidate and uncertainty indicators;
- verification checks, thresholds, and results;
- human reviewer, decision, rationale, and any override;
- downstream action identifier and recovery status; and
- later monitoring results, incidents, and corrective actions.
Evidence is not maximal logging. Sensitive data should not be copied indiscriminately, and retention should follow applicable rules. The objective is decision traceability: enough to test whether the control operated and whether the outcome remained acceptable.
Open questions. What evidence is authoritative? Which failures must stop the workflow rather than create a warning? Who may override, and is override behavior monitored? How quickly can an action be reversed? What performance or incident threshold forces revalidation? Who can suspend the system when its business owner disagrees?
The doctrine succeeds when it avoids two symmetrical errors: treating AI output as a decision, and treating human presence as assurance. AI should generate where variation is useful. Systems should verify what can be made explicit and repeatable. Humans should decide where consequences require accountable judgment. The control intensity should follow the risk—not the enthusiasm surrounding the technology.
Source register
- NIST, Artificial Intelligence Risk Management Framework Core (AI RMF 1.0, 2023)
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1, July 2024; page updated April 2026)
- U.S. Government Accountability Office, Artificial Intelligence: An Accountability Framework for Federal Agencies and Other Entities (GAO-21-519SP, June 2021)
- Federal Reserve SR 26-2 transmitting the Federal Reserve, OCC, and FDIC Revised Guidance on Model Risk Management (April 17, 2026)
- The Institute of Internal Auditors, Artificial Intelligence Auditing Framework, 2nd Edition (September 2024)
Corrections
No corrections recorded.
Report a correction via LinkedIn with the note title and the supporting source.