When an AI assistant hands you a number, a narrative, or a recommendation — what should you do with it before you act? A governance-grade companion to the AI Analysis Audit Guidebook.
Every existing framework for AI in the enterprise — NIST AI RMF, ISO 42001, the Big Four governance playbooks — addresses how to audit AI systems. None of them give a senior leader a tactical protocol for the moment that actually matters: you are thirty minutes from a decision, you have AI-generated analysis in front of you, and you need to know whether to trust it.
AI-assisted analysis is not a deliverable. It is the residue of a conversation, against data that refreshes, with assumptions no one has written down. The traditional audit chain assumed a human analyst sat between the data and the decision-maker and owned the reconciliation — that layer has collapsed. The C-suite discipline for AI is not trust the model or distrust the model; it is knowing, for each output, how much weight the evidence actually supports. This document gives you the tools.
AI reports what is in the file, not what is true in the business. Every executive action on an AI output is a bet that the two are the same.From the AI Analysis Audit Guidebook — Principle Two
| Tool | What it answers | Time to use |
|---|---|---|
| Trust Equation | How much scrutiny does this specific output require before I act? | 2 minutes |
| Trust-Tier Assessor | Interactive: what tier is this decision, and what does that mean operationally? | 3 minutes |
| Five Questions | The pre-action checklist an executive can run before a meeting. | 5 minutes |
| Four Roles | Which hat am I wearing right now, and what does it require of me? | Situational |
| Confidence Gap Check | Is the AI's narrative more confident than the numbers actually support? | 3 minutes |
| Second-Opinion Protocol | The AI-era equivalent of a concurring partner review. | 10 minutes |
| Escalation Matrix | Who gets called, when, and what gets paused — per failure mode. | Reference |
Read sections 3, 4, and 6. The Trust Equation tells you how much to worry. The Assessor turns that into a concrete action path for the output on your desk. The Five Questions is the pre-meeting checklist. Everything else is depth — valuable, but sections 3, 4, and 6 are what you actually run on a Tuesday afternoon before a board call.
A scan of the current landscape — NIST AI RMF, ISO 42001, the AICPA/CAQ guidance, PwC's Responsible AI audits materials, Deloitte's Trustworthy AI framework, the WEF AI C-Suite Toolkit, the IIA's AI Auditing Framework — reveals a consistent pattern. The published work concentrates on governing AI systems: how models are built, validated, deployed, monitored, and documented. Excellent guidance. Wrong problem for the question this document addresses.
The question this document addresses is different: a C-suite leader has received an AI-generated analytical output. What do they do with it between the moment of receipt and the moment of action? That window is where the damage happens, and it is effectively undocumented.
| Reason | How it manifests | Why existing frameworks don't fix it |
|---|---|---|
| The analyst-in-the-middle evaporated | Executives now work with AI directly. The junior analyst who used to reconcile numbers before they reached the CFO is gone, or has been reassigned to higher-order work. | Governance frameworks assume a development/deployment lifecycle. They do not address ad hoc use of general-purpose AI by senior leaders. |
| The artifact is a conversation, not a document | An AI output is the end-state of a multi-turn session with drifting assumptions, data that refreshed, and prompts no one kept. | Audit frameworks assume static artifacts with clear provenance. AI conversations have neither by default. |
| No one owns the executive layer | Internal audit audits the system. The CDO/CIO governs the platform. The analyst audits the data. The executive audits… nothing, formally. | There is no profession-level standard for what a decision-maker owes the decision when AI produced the evidence. |
Traditional audit standards are built by regulators, standard-setters, and academics, then adopted by practice. For AI-output audit, that sequence is too slow. The discipline has to be built from the user side — by practitioners who feel the gap — and then codified later. This document is an attempt to do the first part well enough to make the second part possible.
This framework does not arrive on empty ground, and it is stronger for saying so plainly.
The question of how much to trust a machine's output — and when that trust is miscalibrated — is a mature research field. Work on trust calibration and “appropriate reliance” (Lee & See, 2004) and on automation bias and algorithm aversion (Parasuraman & Riley, 1997; Dietvorst et al., 2015) established decades ago that people both over- and under-rely on automated advice, and that the failure is rarely the tool. A newer literature extends this to generative AI specifically, naming the way fluent, confident output earns trust it hasn’t earned on the merits.
At the organizational layer, the ground is well covered — NIST’s AI RMF, ISO 42001, and Gartner’s AI TRiSM all govern how AI systems are built, validated, and monitored. And at the decision layer, some executive-facing guidance now sorts decision types by impact and reversibility, mapping which categories a human must own.
What none of these provide is the thing a leader actually needs at 4:40 on a Tuesday: a protocol for a single output, before a single action. The research describes the phenomenon but does not hand the decision-maker a procedure. The governance frameworks operate a layer above the moment of action. The decision-category maps use reversibility and stakes — but stop short of the axis that determines whether the output can be checked at all: auditability, the traceability of every number back to a source a reviewer can follow. That third axis, drawn from the audit tradition rather than the AI-governance one, is this framework’s contribution — and the reason it resolves not to a category but to a concrete, per-output next step.
Three variables determine how much scrutiny an AI output needs. They combine multiplicatively — a weakness on any axis is not offset by strength on the other two.
If the decision turns out to be wrong, how hard is it to unwind?
What is the magnitude of harm if the output is materially wrong?
Can the output be traced back to source data with formulas a reviewer can follow?
The intuition is not difficult but is easy to get wrong in the moment. A low-reversibility, high-stakes, low-auditability output is the maximum-danger configuration — and it is the configuration most likely to appear when a senior executive pastes a P&L into an AI tool and asks for commentary twenty minutes before a meeting. The fact that the stakes are high does not make the output trustworthy; the fact that the AI sounded confident does not make it auditable; the fact that you are senior does not make the decision reversible.
The equation collapses to a four-tier classification.
| Tier | Configuration | What it means in practice |
|---|---|---|
| T1 — Directional | High reversibility, low-medium stakes, any auditability | Use as input. Verify numbers spot-check only. Act. |
| T2 — Decision-Support | Medium reversibility OR medium stakes, medium+ auditability | Require aggregate reconciliation. Name a reviewer. Act with note in file. |
| T3 — Board-Grade | Low reversibility, high stakes, medium+ auditability required | Tier-1 or Tier-2 audit mandatory (see AI Analysis Audit Guidebook). Named accountable executive. Second opinion. Do not act without sign-off. |
| T4 — No-Go | Low auditability regardless of other variables, OR an irreversible decision class (below) | AI output is reference only. Human analysis is the deliverable. Do not let AI narrative drive the call. |
Executives often treat confidence in the AI's tone as a proxy for auditability. It is not. Modern AI produces fluent, structured, sourced-sounding output regardless of whether the underlying numbers are correct. Tone is evidence of writing capability, not of analytical accuracy. Keep the two separated.
Answer the three questions below for the specific AI output on your desk. The assessor classifies it into a tier and prescribes next steps.
A C-level leader interacts with AI output in four distinct roles. The obligations differ. Confusing them is the source of most over- and under-trust errors.
Most AI-related executive errors occur at role transitions. A leader runs a quick Operator task, hands the output to a colleague (who is now a Consumer), and neither records the assumptions that were made. When a Sponsor-level question surfaces six weeks later — "why did we exclude off-rent days?" — the answer has been lost. Name the role explicitly at the top of any AI-assisted workstream.
Run this list before any Tier-2 or Tier-3 action. It takes five minutes. Check each off.
The checkboxes above are interactive — the state does not persist, so if you want to use this as a real-time tool, print the page or capture the state as you go. A fuller Power Automate / SharePoint integration is out of scope for v1 but should be trivial to layer on.
The single most dangerous pattern in AI-assisted executive work: the narrative is more confident than the underlying numbers support. The AI writes with the authority of a consulting memo; the data beneath it has the fragility of a first draft.
| Signature | How it shows up | How to pressure-test it |
|---|---|---|
| Confident direction with fuzzy magnitude | "Revenue is trending meaningfully down." Down by what? Since when? Compared to what baseline? | Ask for the specific number, period, and comparator. If the AI has to go recompute, the original statement was asserted without data. |
| Causal claims from correlational data | "The new pricing strategy caused the utilization drop." With what control? What else changed? | Ask: "What would I need to be true for this claim to be wrong?" If nothing, it is not a claim — it is a narrative. |
| Hedged numbers with unhedged conclusions | Numbers carry "approximately," "may," "estimated." Conclusion says "we should." The caveats did not carry through. | Rewrite the conclusion using the same hedging language as the inputs. If it now reads as weak, the original was over-reaching. |
| Synthesis that exceeds the evidence | The AI combines three weak signals into one strong conclusion. Each signal was tentative; the conclusion is definitive. | Demand a confidence note per input. If any is "low," the synthesis is at most "medium." Arithmetic on confidence is not optimistic. |
Never accept an AI's interpretation until the underlying numbers are audited. The story is only as good as the data it sits on.AI Analysis Audit Guidebook — Common Failure Modes
In financial audit, a concurring partner reviews the engagement partner's work before sign-off. It is the single highest-ROI control in the audit standard. The AI equivalent exists, is cheap, and is almost never used. It takes two forms.
Give the same prompt and data to a second AI from a different provider. Different training data, different architectures, different biases. If the two agree within tolerance, the shared answer is more credible. If they diverge, the divergence itself is evidence — you now have a specific disagreement to investigate rather than a single confident answer to trust or reject blindly.
Do not share the first AI's output with the second AI. That defeats the independence. Give each the raw source and the raw question. Compare the outputs. This is the same reason auditors keep the concurring partner walled off from the engagement team's conclusions.
Same model, new conversation. No history. Same prompt, same data. If the result materially differs, you have evidence of multi-turn drift in the original session — the AI's assumptions had shifted over the course of the conversation, and the final output reflected cumulative drift rather than the question you started with.
| Protocol | When to use | Time / cost |
|---|---|---|
| Form A — Cross-model | Tier 3 decisions. High-stakes strategic analysis. Anything going to the board. | 10–15 minutes. Cost of a second AI subscription. |
| Form B — Fresh-session | Tier 2+ work produced in a long conversation (>20 turns, or the session spans days). | 5–10 minutes. No additional cost. |
| Combined | Irreversible decisions. External disclosures. Capital allocation. | 30 minutes. Cheap relative to the decision. |
If two models agree, that raises confidence but does not guarantee correctness — they may share a failure mode (e.g., both were trained on the internet; both have been taught the same wrong fact). Cross-model agreement plus a source-data trace is the combination that should give you actual comfort. Neither alone is sufficient.
The AI Analysis Audit Guidebook identifies failure modes. This section maps each to an escalation path: who gets called, when the decision pauses, and what recovers.
| Failure mode detected | What it implies | Who gets notified | What pauses |
|---|---|---|---|
| Static values instead of formulas | Output cannot be audited without rebuilding. Decision evidence chain is weak. | Analyst who produced it; reviewer named in workstream. | Any Tier-2 or Tier-3 action. Rebuild with formulas before proceeding. |
| Silent filter or scope drift | Output describes a different question than was asked. Conclusion may be valid, but for the wrong population. | Analyst; prompt author if different. | All use of the output. Re-run with explicit scope statement in the prompt. |
| Source data completeness gap | The file is not what the system of record shows. Numbers are arithmetically correct but describe incomplete reality. | Data owner; CIO or CDO for repeated patterns; internal audit for material gaps. | Decision pauses. Reconcile sample to ERP/GL before release. |
| Multi-turn drift | The analysis in hand is not what was asked for at session start. Late-session conclusions inherit accumulated drift. | Analyst; executive consumer if already distributed. | All downstream use. Fresh-session re-run per Protocol Form B. |
| Confidence gap (narrative exceeds data) | Conclusion is more definitive than evidence supports. High public- or board-harm risk. | Analyst; communications lead for external-facing content; legal for disclosure context. | Any external communication. Rewrite with hedging matched to inputs. |
| Unreconstructable prompt history | Decisions are traceable only to "the AI said so." Cannot defend under scrutiny six months later. | Analyst; whoever owns the AI governance policy (general counsel, chief risk officer, internal audit). | Irreversible decisions. Rebuild with decision log per AI Analysis Audit Guidebook §Prompt Archaeology. |
Escalation matrices that do not include "what pauses" are aspirational. The whole value of this framework is the willingness — culturally and procedurally — to pause a decision when the evidence is inadequate. A senior leader who pauses for 24 hours to audit a questionable AI output will be right more often than one who ships on the AI's timeline.
For organizations subject to SOX, operating under COSO, or with a mature internal audit function, the executive-layer audit discipline fits neatly into existing structures. The integration is conceptual, not structural — no new committees are required.
| Existing framework | How AI-output audit maps in | What to add to policy language |
|---|---|---|
| COSO Internal Control | AI-output audit is a detective control under Control Activities (Principle 10) and sits within Risk Assessment (Principles 6–9) when AI touches financial-reporting decisions. | "Any AI-assisted analysis used in decisions with financial-reporting impact must be classified per the Trust-Tier framework and audited to the corresponding tier before use." |
| SOX §404 | AI used in the production of SOX-relevant numbers (estimates, journal entries, close-process analytics) is in scope for controls testing. | SOX narratives must explicitly flag where AI produces or influences a reported number, and document the audit tier applied. |
| Internal Audit Charter | Internal audit is well-positioned to own the Tier-3 review function — it already has the independence and the audit tradecraft. | Expand charter to include "review of AI-assisted executive analytical output for irreversible decisions." |
| Board Reporting | Board materials are the highest-stakes context in which AI output appears. They should carry provenance tags at minimum. | "All board materials derived in part from AI analysis must cite the audit tier applied and name the accountable executive." |
| NIST AI RMF | NIST covers the system; this framework covers the output consumption layer. Complementary, not duplicative. | Adopt NIST for system governance, adopt this framework for user-side output audit. Same policy, two layers. |
If your company does not have a formal internal audit function or SOX program, the Five Questions and the Trust-Tier Assessor are themselves the program. Write them into your leadership handbook, require them above a stakes threshold, and keep a simple decision log. That is enough to materially outperform most peers, who have no user-side AI audit discipline at all.
AI is a production layer, not a review layer — “the AI reviewed it” is not a control. The human who acts on an output owns it, and “the AI did it” is not a defense but an admission that no one reviewed it. Where an output crosses multiple desks, ownership is joint and non-transferable: each reviewer owns what their level is positioned to catch, and “I assumed someone upstream checked” is itself the review failure. Because fluent, confident output suppresses the doubt that used to trigger scrutiny, every layer must perform an active, documented review. This is Principle 7 of the companion Audit Guidebook — and the reason the review skill matters more in the AI era, not less.
Some decisions should not be driven by AI analysis regardless of how good the audit is. They involve either reversibility characteristics or evidentiary standards that AI output cannot satisfy as a class.
| Decision class | Why AI output is insufficient | What AI can still contribute |
|---|---|---|
| Termination of an employee | Due process, legal exposure, dignity. The record must be built by humans who can testify to what they saw. | Summarizing prior documentation; drafting neutral language. Not the decision. |
| Public disclosure of forward-looking guidance | Securities law standards for basis and reasonableness are human-judgment standards. AI cannot be a witness. | Scenario modeling; sensitivity analysis to stress-test human judgment. |
| Going-concern and material-estimate judgments | GAAP/IFRS require documented management judgment. AI as the decision-maker violates the standard directly. | Data preparation; consistency checks across periods. |
| Irreversible capital allocation without human reconciliation | The cost of being wrong is permanent. Audit discipline requires a named human who agrees, in writing. | Alternatives generation; cross-check against base rates. |
| Safety-critical operational decisions | Lives or serious injury at stake. AI models are probabilistic; safety standards are not. | Anomaly detection, early warning — with a human in the loop for action. |
| Board-governance votes | Fiduciary duty is personal. Delegation to AI is a breach as a category, not as a matter of quality. | Pre-read preparation; comparison of vendor proposals to prior precedent. |
Observe that none of these prohibit AI assistance — they prohibit AI authorship of the decision. The line is drawn at who is accountable, not at who did the typing.
For practitioners, analysts, and aspiring senior leaders who want to operationalize this framework as a personal capability rather than read it as a reference.
| Week | Practice | Artifact to produce |
|---|---|---|
| Week 1 — See | Apply the Trust-Tier Assessor to every AI output you produce or receive. Do not try to change behavior yet. Just classify. | A running list of tier classifications, with actual decisions made for each. At week's end, review which you acted on without the corresponding audit level. |
| Week 2 — Audit | Pick the highest-stakes AI output from Week 1. Build a real Tier-1 audit of it per the AI Analysis Audit Guidebook methodology. | A worked audit sheet. Either the numbers tie, or they do not. If they do, note the confidence that adds. If they do not, note what you almost acted on. |
| Week 3 — Prompt | For each new AI task, write the prompt to include explicit scope, exclusion criteria, and a request for auditable output. | A personal library of five prompt templates, one per common task type, in the pattern from the AI Analysis Audit Guidebook §Audit Prompts. |
| Week 4 — Review | Run the Second-Opinion Protocol on the highest-stakes Tier-2 output of the week. Compare. Document the divergence or the agreement. | A short memo to yourself: what the divergence was, what it taught you about the original output, and whether the decision stands. |
The skill described in this document is not technical. It is a form of professional skepticism — the same habit that has distinguished strong auditors and strong executives for a century, now applied to a new evidentiary medium. The AI will change. The models will improve. The obligation to know what you are acting on, and why, will not.