AAgentProof

Workflow module

The agent test pack

The seven checks that turn 'we tested it' into a documented record — where a test never run is a gap, not a pass.

Seven checksNot-run is a gapAll-seven-pass seamNo verdict

The seven checks that turn 'we tested it' into a documented record — where a test never run is a gap, not a pass.

Why it matters

Agents fail in ways ordinary software does not. They hallucinate, get talked out of their instructions, leak data, and reach past their tool boundaries. 'We tested it' stays hopeful until it becomes a record of what was actually checked and what was found.

What good looks like

Seven agent-relevant checks — Hallucination, Prompt injection, Jailbreak, Data leakage, Tool boundary, Unapproved action, and Bias — each recorded with a real outcome, a method, a summary, severity where relevant, a tester, and a date. A check that has never been run shows Not run, a visible gap, rather than silently reading as a pass or a fail. Unknown is never treated as no.

What can go wrong

Testing is a vibe rather than a record. A check that was never run is assumed to be fine. One never-run or unresolved check is hidden inside an overall green light. Results arrive with no method and no date, so no one can tell what was actually exercised.

What AgentProof checks

AgentProof records the seven-type machine with a first-class Not-run state and a three-key outcome (Pass, Fail, Not run), and surfaces per-type coverage and gaps. The one readiness seam is deliberately conservative: every one of the seven latest-run checks must be Pass for tested-before-deployment to resolve — any Fail or any Not run leaves it unresolved, and never reads as no. It documents behaviour, never a whole-agent verdict.

Key terms

The exact vocabulary this part of the record uses — grounded in the shipped product model.

Test types
hallucination, prompt_injection, jailbreak, data_leakage, tool_boundary, unapproved_action, bias.
Outcomes
pass, fail, not_run — a check never run is a visible gap.
Severity
low, medium, high, critical — recorded where relevant.
Tested-before-deployment seam
Resolves only when all seven latest-run checks are Pass; any fail or not-run leaves it unresolved.

Keep reading

Build this record for your own agents

AgentProof turns each part of this into one documented, evidence-backed record — before your compliance review.

Ready to build a compliance-readiness record for your own agent?

Request a compliance readiness pilot to apply this guidance to a real agent.

AgentProof builds a compliance-readiness record, not an official audit, and it does not speak on behalf of any vendor.