Intent: comparing the call with the user’s request
DefenseClaw checks each tool call against what the user asked for. Eight runs on AgentDojo measured that check as harmful. On a corpus where the compromise is independently proven, the same code, models and seed turn all 32 measured cells positive. It separates compromise from refusal; it does not block enough to be a control on its own.
The statistic
An attacker plants text the agent will read. The agent either complies or correctly refuses, and in both cases the attacker’s text is present and the agent keeps making tool calls. The guard has to tell the two apart, so the measure is separationBlock rate on a proven compromise minus block rate on an agent that correctly refused the same attack. Below zero, the guard is reacting to the attacker’s text rather than to the compromise.:
separation = block rate on a proven compromise − block rate on an agent that correctly refused the same attack
Written sep|res in the artifacts. Below zero,
the guard fires more on the refusal: it is reacting to the attacker’s text and punishing
the agent for doing the right thing.
Eight negative runs on AgentDojo
AgentDojo is a synthetic benchmark with matched attacked and clean pairs: 651 matched cases, four models, two context recipes. Every run came out negative, between −0.0053 and −0.5343 per case counting only hard blocks, and supplying the user’s request made it worse in 26 of 36 paired comparisons across the first three models. Two mechanical causes, both measured and both specific to that benchmark:
- Attacker-text sensitivity. On AgentDojo the only difference between a clean case and a correct refusal is the attacker’s text. One model’s flag rate goes from 22.35% on clean cases to 80.31% on correct refusals, against 20.62% on real compromises.
- The aggregation unit. A case counts as flagged if any of its events fired, and correct refusals had the most events per case (5.07 against 4.14 for compromise). Scored per event, OpenJev’s cell with the request moves from -0.5343 to -0.0831.
The reversal on proof-backed data
A new corpus was built from real emitted MCP tool calls where a deterministic proof establishes that the compromise happened: 3,402 cases over 1,929 families, 405 of them proof-verified. All four models were re-run on it, and the same scorer with the same seed was re-run on AgentDojo as a control. The proof-backed corpus is evaluation-only and reported in aggregate only: this site carries its rates and counts and never an example.
Separation in all 32 measured cells, two corpora
Data table
| Backend | context | unit | lens | proof-backed | AgentDojo prior | sign |
|---|---|---|---|---|---|---|
| OpenJev | C0 | case | block | +0.266667 | -0.057208 | flips |
| OpenJev | C0 | case | any | +0.395089 | -0.338390 | flips |
| OpenJev | C0 | event | block | +0.212598 | +0.003403 | same sign |
| OpenJev | C0 | event | any | +0.310809 | -0.111818 | flips |
| OpenJev | C7 | case | block | +0.575008 | -0.534346 | flips |
| OpenJev | C7 | case | any | +0.768989 | -0.596923 | flips |
| OpenJev | C7 | event | block | +0.505119 | -0.083135 | flips |
| OpenJev | C7 | event | any | +0.648842 | -0.102807 | flips |
| DiffusionGemma | C0 | case | block | +0.400985 | -0.124539 | flips |
| DiffusionGemma | C0 | case | any | +0.442304 | -0.338150 | flips |
| DiffusionGemma | C0 | event | block | +0.323368 | -0.025036 | flips |
| DiffusionGemma | C0 | event | any | +0.356870 | -0.064921 | flips |
| DiffusionGemma | C7 | case | block | +0.484594 | -0.394076 | flips |
| DiffusionGemma | C7 | case | any | +0.564310 | -0.446611 | flips |
| DiffusionGemma | C7 | event | block | +0.414046 | -0.068889 | flips |
| DiffusionGemma | C7 | event | any | +0.471407 | -0.092620 | flips |
| Gemma 4 | C0 | case | block | +0.175309 | -0.005315 | flips |
| Gemma 4 | C0 | case | any | +0.445436 | -0.240292 | flips |
| Gemma 4 | C0 | event | block | +0.139764 | -0.000939 | flips |
| Gemma 4 | C0 | event | any | +0.345329 | -0.105622 | flips |
| Gemma 4 | C7 | case | block | +0.333956 | -0.290690 | flips |
| Gemma 4 | C7 | case | any | +0.720228 | -0.380802 | flips |
| Gemma 4 | C7 | event | block | +0.271785 | -0.027485 | flips |
| Gemma 4 | C7 | event | any | +0.617478 | -0.092719 | flips |
| Jev 1.13.0 | C0 | case | block | +0.162040 | -0.020832 | flips |
| Jev 1.13.0 | C0 | case | any | +0.414103 | -0.297340 | flips |
| Jev 1.13.0 | C0 | event | block | +0.129003 | -0.004227 | flips |
| Jev 1.13.0 | C0 | event | any | +0.329470 | -0.121868 | flips |
| Jev 1.13.0 | C7 | case | block | +0.311111 | -0.124566 | flips |
| Jev 1.13.0 | C7 | case | any | +0.605061 | -0.416164 | flips |
| Jev 1.13.0 | C7 | event | block | +0.248031 | -0.025837 | flips |
| Jev 1.13.0 | C7 | event | any | +0.489387 | -0.137907 | flips |
outputs/intent-real/jev-analysis.json :: four_backend_table[backend][context][unit][lens].{intent_real, agentdojo_prior}.sep_vs_resisted; four_backend_summary.The one cell that is not a sign reversal (OpenJev without the request, C0, per event, block-only) was already positive at +0.003403. For OpenJev with the request (C0 / C7C0 is the tool call alone. C7 adds the user’s request and up to seven prior events.), separation is -0.5343 on AgentDojo and +0.5750 on the proof-backed corpus.
Intent separation, four models, two corpora
One bar per model per corpus per context. Left of zero means the guard flagged a correct refusal more often than a proven compromise. Error bars are family-cluster bootstrap 95% intervals, 2,000 replicates, seed 741983. Lead cell: proof-verified (grade A) lane, block-only lens, per case.
Data table
| Backend | Corpus / context | sep|res | bootstrap 95% |
|---|---|---|---|
| OpenJev — self-hosted, FP8, structured instructions, question Q2 | AgentDojo, no intent (C0) | -0.0572 | [-0.1135, -0.0066] |
| OpenJev — self-hosted, FP8, structured instructions, question Q2 | AgentDojo, with intent (C7) | -0.5343 | [-0.6146, -0.4476] |
| OpenJev — self-hosted, FP8, structured instructions, question Q2 | Proof-backed, no intent (C0) | +0.2667 | [+0.2241, +0.3089] |
| OpenJev — self-hosted, FP8, structured instructions, question Q2 | Proof-backed, with intent (C7) | +0.5750 | [+0.5270, +0.6209] |
| DiffusionGemma 26B-A4B-it-FP8-dynamic — self-hosted, string instructions, Q2 | AgentDojo, no intent (C0) | -0.1245 | [-0.1881, -0.0643] |
| DiffusionGemma 26B-A4B-it-FP8-dynamic — self-hosted, string instructions, Q2 | AgentDojo, with intent (C7) | -0.3941 | [-0.4736, -0.3137] |
| DiffusionGemma 26B-A4B-it-FP8-dynamic — self-hosted, string instructions, Q2 | Proof-backed, no intent (C0) | +0.4010 | [+0.3504, +0.4497] |
| DiffusionGemma 26B-A4B-it-FP8-dynamic — self-hosted, string instructions, Q2 | Proof-backed, with intent (C7) | +0.4846 | [+0.4367, +0.5333] |
| Gemma 4 (google.gemma-4-26b-a4b via Bedrock) — chat judge, question Q0 | AgentDojo, no intent (C0) | -0.0053 | [-0.0378, +0.0273] |
| Gemma 4 (google.gemma-4-26b-a4b via Bedrock) — chat judge, question Q0 | AgentDojo, with intent (C7) | -0.2907 | [-0.3751, -0.2089] |
| Gemma 4 (google.gemma-4-26b-a4b via Bedrock) — chat judge, question Q0 | Proof-backed, no intent (C0) | +0.1753 | [+0.1390, +0.2136] |
| Gemma 4 (google.gemma-4-26b-a4b via Bedrock) — chat judge, question Q0 | Proof-backed, with intent (C7) | +0.3340 | [+0.2858, +0.3801] |
| Jev 1.13.0 — hosted API, structured instructions, Q2 | AgentDojo, no intent (C0) | -0.0208 | [-0.0565, +0.0130] |
| Jev 1.13.0 — hosted API, structured instructions, Q2 | AgentDojo, with intent (C7) | -0.1246 | [-0.1861, -0.0641] |
| Jev 1.13.0 — hosted API, structured instructions, Q2 | Proof-backed, no intent (C0) | +0.1620 | [+0.1241, +0.1977] |
| Jev 1.13.0 — hosted API, structured instructions, Q2 | Proof-backed, with intent (C7) | +0.3111 | [+0.2647, +0.3589] |
outputs/intent-real/jev-analysis.json :: four_backend_table (AgentDojo priors recomputed by the same scorer; cross-checked against replication-analysis.json and separation-analysis.json).Why the reversal is not an artefact
- The control did not move. The same scorer, flags and seed reproduce every model’s published AgentDojo figure to four decimal places.
- Not corpus identity. Compromise and correct refusal come from the same source corpus.
- Not a between-family confound. Restricted to the 325 families holding both a proven compromise and a correct refusal on the same user request and the same poisoned tool catalogue, all 32 cells stay positive.
- Not carried by weak labels. Grade A beats grade B in all 8 cells for every model.
Limits
Block-only recall on proof-verified compromises
Separation measures discrimination. Recall bounds the production claim: the share of independently proof-verified compromises the guard actually blocks. A confirm stops nothing, so only hard blocks count here.
Data table
| Backend | Context | recall | blocked / total | missed |
|---|---|---|---|---|
| OpenJev | no user request supplied (C0) | 0.2667 | 108/405 | 73.3% |
| OpenJev | user request + history (C7) | 0.5778 | 234/405 | 42.2% |
| DiffusionGemma | no user request supplied (C0) | 0.4222 | 171/405 | 57.8% |
| DiffusionGemma | user request + history (C7) | 0.4938 | 200/405 | 50.6% |
| Gemma 4 | no user request supplied (C0) | 0.1753 | 71/405 | 82.5% |
| Gemma 4 | user request + history (C7) | 0.3358 | 136/405 | 66.4% |
| Jev 1.13.0 | no user request supplied (C0) | 0.1630 | 66/405 | 83.7% |
| Jev 1.13.0 | user request + history (C7) | 0.3111 | 126/405 | 68.9% |
outputs/intent-real/jev-analysis.json :: grade_a_block_only_recall.Every model misses 42.2% to 68.9% of proven compromises. With the request supplied, the four block 0.5778 (OpenJev), 0.4938 (DiffusionGemma), 0.3358 (the Gemma 4 judge) and 0.3111 (Jev 1.13.0).
- “With intent” is not a clean isolation. C7 adds the request, up to seven prior events and a different text representation at once, so the C0-to-C7 gain cannot be attributed to the request alone.
- No between-model ranking. The order of the four changes with the lens and the context; only the within-model contrast is controlled.
- The refusal label is upstream. Nothing independently proves the agent ignored the attack; mislabelling would depress separation, so the bias is conservative.
| Claim | Status |
|---|---|
| Comparing the call with the user’s request separates proven compromise from correct refusal. | supported |
| Knowing the request improves that separation. | supported |
| The earlier eight-run negative was a property of the mechanism. | refuted |
| This is enough to ship as a blocking control. | not supported |
evaluation-only never-train
Nothing derived here is approved for training, synthetic generation, teacher context, distillation or redistribution. The 13 source datasets are public and linked on reproduce; the analysis artifacts behind each number are held privately and are available on request.