AI agent assurance
50 requirements in six domains that the internal auditor uses to test AI agents: those of the business and those of the audit function itself. Under whose authorisation did the agent act, and where is the evidence that it stayed within that mandate?
The problem
From automation claim to demonstrable control
Vendors of audit AI agents emphasise automation and efficiency in their product information, far less how mandate, authorisation, error measurement and human intervention are demonstrably controlled.
Sales claims, no tested control
Diligent, AuditBoard (now Optro) and TeamMate claim automation of up to 87 percent of SOX tasks, without making concrete how mandate and authorisation are controlled.
One framework tests every agent
Fixed requirements instead of a questionnaire reinvented per project, applied to the vendor’s agent, the business agent and the audit function’s own agent.
Identity and mandate left unresolved
NIST and Singapore’s IMDA framework flag the same question: how does an agent identify itself, how does it prove its authority, and on whose behalf did it act.
Own technical identity per agent
Every agent authenticates with its own technical identity, not under a shared service account, with the mandate traceable back to the authorising human.
Kill switch and error rate unknown
The stop mechanism has never been tested, and as long as failed and corrected actions are not recorded, any statement about reliability is a guess.
Tested, measured and documented
The kill switch has been genuinely exercised at least once with the result recorded, and the error rate is measured on a representative set of tasks.
This is what it looks like
A selection from AI agent assurance
Click a thumbnail for that screen, or the image itself for the full-size view.
The framework in numbers
Fixed requirements, not a loose questionnaire
6 domains
From inventory and identity to compliance and reporting.
50 requirements
Spread across the six domains, each separately testable.
23 mappings
To the shared Unified Controls, each mapping assessed individually on its merits.
27 requirements
Deliberately unmapped: the Unified Controls are ISO 27001-based and have no fair counterpart for, say, an exercised kill switch.
What the framework offers
The reverse question: we test the agent
A fixed set of requirements instead of a questionnaire reinvented for every engagement, and without the sales pitch about what the agent automates.
Testing instead of selling
The framework does not assess what an agent promises to automate, but whether its actions are demonstrably authorised, bounded, controlled and recorded.
Applies to the audit function itself
The same requirements apply to the agent handling HR processes and to the agent searching audit files. An audit function tests itself first this way.
Data minimisation, including data riding along
Five requirements test whether the agent sees no more than its task requires, and whether personal data riding along in attachments, exports and mailboxes is recognised.
AI Act assessed per agent
Article 50 applies to agents that interact directly with people, Article 4 requires AI literacy; domain F determines this per use case, not with one rule for all agents.
23 substantively validated mappings
Every mapping to a Unified Control was assessed individually on its merits.
Part of the CRAFT library
The framework sits in CRAFT alongside ISO/IEC 42001, NIST AI RMF and OWASP LLM, and evolves with the rest of the library.
The six domains
How the framework covers the agent, from inventory to board reporting
The framework works through the agent from the outside in: first what is running, then under what mandate, then who can intervene, what is evidenced, how reliable the behaviour is, and what has to be accounted for.
Inventory, identity and mandate
Which agents are running, under which own technical identity, and under whose authorisation: which data may they access and is that boundary technically enforced.
Human intervention
Where is prior approval required, is that a real assessment or routine click-through, who can stop the agent, and has that kill switch ever been exercised.
Evidence and traceability
An action log the agent cannot alter itself, traceability to the instruction and model version, and a retention period derived from the decision’s limitation period.
Reliability and testing
Was the agent tested beforehand against a set standard, is there a measured error rate, and is silent degradation noticed.
Compliance and reporting
Disclosure to whoever the agent communicates with, AI literacy for those who steer it, accountability to the board and audit committee, and the regime for purchased agents.
Who it is for, and in the suite
What the testing delivers
The testing gives the CAE a substantiated answer to the question the board, external auditor or regulator will ask sooner or later.
The CAE: which agents are running, who authorised them, which personal data they touched, and where the evidence is that they stayed within that authorisation.
The audit function itself: the same requirements apply to the audit function’s own agent, which searches audit files containing personal data.
Privacy and compliance: five requirements specifically test data minimisation and personal data riding along in attachments, exports and mailboxes.
The framework in the suite
Test the agents already running
Browse the full framework in the CRAFT library, or let us show you in a demonstration what the testing would deliver for your agents.