Chatbot safety measured refusal. Agent safety adds a second axis: after the jailbreak, does the model still KNOW HOW to be harmful? 110 malicious multi-step tasks across 11 categories measure both.
The chatbot-era metric that agent misuse outgrew.
A harmful agent task has two independent failure modes for safety: the model refuses (good) or complies (bad); and if it complies, does it retain its capabilities (multi-step planning, tool orchestration) under the jailbreak? AgentHarm scores both. The tasks are deliberately explicit — unambiguously malicious, no ambiguity to hide behind — spanning fraud, cybercrime, harassment, and 8 more categories. Each ships with a rubric: refusal correctness, step completeness, and tool-use quality. The rubric design makes the benchmark extensible (augmentations multiply coverage without hand-writing new tasks).
The misuse question chatbot evals never asked.
Refusal evals are the bouncer: can the model turn away trouble at the door? AgentHarm adds the black-belt test: if trouble gets in, can the model actually fight — plan the sequence, use the tools, finish the multi-step job? A jailbroken model that passes both black-belt and compliance isn't a clumsy drunk — it's a hired problem. The benchmark measures exactly that severity.
What a malicious agent task contains — the structure that enables two-axis scoring.
What two-axis measurement revealed about 2024 agents.
The asymmetry that defines agent misuse risk.
| Safety axis | What it measures | Failure mode |
|---|---|---|
| Refusal | does the model decline? | compliance with malicious intent |
| Capability retention | does the jailbroken agent still work? | incompetent compliance — the lucky case |
| Both (AgentHarm) | refusal + post-jailbreak capability | competent compliance — the dangerous case |
The two-axis design: 'it complied but fumbled' and 'it refused' are different safety outcomes — and only the second is acceptable.
Agent safety evaluation adopted the two-axis contract.
Check your understanding of the key concepts from AgentHarm.
Everything you need to remember about this paper.