Workflow module
The agent test pack
The seven checks that turn 'we tested it' into a documented record — where a test never run is a gap, not a pass.
The seven checks that turn 'we tested it' into a documented record — where a test never run is a gap, not a pass.
Why it matters
Agents fail in ways ordinary software does not. They hallucinate, get talked out of their instructions, leak data, and reach past their tool boundaries. 'We tested it' stays hopeful until it becomes a record of what was actually checked and what was found.
What good looks like
Seven agent-relevant checks — Hallucination, Prompt injection, Jailbreak, Data leakage, Tool boundary, Unapproved action, and Bias — each recorded with a real outcome, a method, a summary, severity where relevant, a tester, and a date. A check that has never been run shows Not run, a visible gap, rather than silently reading as a pass or a fail. Unknown is never treated as no.
What can go wrong
Testing is a vibe rather than a record. A check that was never run is assumed to be fine. One never-run or unresolved check is hidden inside an overall green light. Results arrive with no method and no date, so no one can tell what was actually exercised.
What AgentProof checks
AgentProof records the seven-type machine with a first-class Not-run state and a three-key outcome (Pass, Fail, Not run), and surfaces per-type coverage and gaps. The one readiness seam is deliberately conservative: every one of the seven latest-run checks must be Pass for tested-before-deployment to resolve — any Fail or any Not run leaves it unresolved, and never reads as no. It documents behaviour, never a whole-agent verdict.
Key terms
The exact vocabulary this part of the record uses — grounded in the shipped product model.
- Test types
- hallucination, prompt_injection, jailbreak, data_leakage, tool_boundary, unapproved_action, bias.
- Outcomes
- pass, fail, not_run — a check never run is a visible gap.
- Severity
- low, medium, high, critical — recorded where relevant.
- Tested-before-deployment seam
- Resolves only when all seven latest-run checks are Pass; any fail or not-run leaves it unresolved.
Keep reading
Build this record for your own agents
AgentProof turns each part of this into one documented, evidence-backed record — before your compliance review.
Ready to build a compliance-readiness record for your own agent?
Request a compliance readiness pilot to apply this guidance to a real agent.
AgentProof builds a compliance-readiness record, not an official audit, and it does not speak on behalf of any vendor.