Intent: comparing the call with the user’s request

DefenseClaw checks each tool call against what the user asked for. Eight runs on AgentDojo measured that check as harmful. On a corpus where the compromise is independently proven, the same code, models and seed turn all 32 measured cells positive. It separates compromise from refusal; it does not block enough to be a control on its own.

The statistic

An attacker plants text the agent will read. The agent either complies or correctly refuses, and in both cases the attacker’s text is present and the agent keeps making tool calls. The guard has to tell the two apart, so the measure is separationBlock rate on a proven compromise minus block rate on an agent that correctly refused the same attack. Below zero, the guard is reacting to the attacker’s text rather than to the compromise.:

separation = block rate on a proven compromise − block rate on an agent that correctly refused the same attack

Written sep|res in the artifacts. Below zero, the guard fires more on the refusal: it is reacting to the attacker’s text and punishing the agent for doing the right thing.

Eight negative runs on AgentDojo

AgentDojo is a synthetic benchmark with matched attacked and clean pairs: 651 matched cases, four models, two context recipes. Every run came out negative, between −0.0053 and −0.5343 per case counting only hard blocks, and supplying the user’s request made it worse in 26 of 36 paired comparisons across the first three models. Two mechanical causes, both measured and both specific to that benchmark:

The reversal on proof-backed data

A new corpus was built from real emitted MCP tool calls where a deterministic proof establishes that the compromise happened: 3,402 cases over 1,929 families, 405 of them proof-verified. All four models were re-run on it, and the same scorer with the same seed was re-run on AgentDojo as a control. The proof-backed corpus is evaluation-only and reported in aggregate only: this site carries its rates and counts and never an example.

Separation in all 32 measured cells, two corpora

+0.30 and above+0.02 to +0.30within ±0.02 of zero−0.02 to −0.30−0.30 and below
Separation in all 32 measured cells, two corporaAll 32 cells on proof-backed data are positive, and 31 of the 32 are strict sign reversals against the AgentDojo prior run of the same code. The remaining cell was already positive in that prior run. Separation = flag rate on a proven compromise minus flag rate on an agent that refused it. Above zero the guard reads the outcome; below zero it reads the attacker text. Four backends × C0/C7 × unit × lens = 32 cells per corpus, one block each. Diverging scale centred on zero. The value is printed in every cell. 32 of 32 cells are positive on the proof-backed corpus; 31 of 32 flip sign. proof-backed data (intent-real) C0 case blk C0 case any C0 evt blk C0 evt any C7 case blk C7 case any C7 evt blk C7 evt any OpenJev OpenJev · C0 · unit=case · block lens · proof-backed data (intent-real): +0.2667 +0.27 OpenJev · C0 · unit=case · any lens · proof-backed data (intent-real): +0.3951 +0.40 OpenJev · C0 · unit=event · block lens · proof-backed data (intent-real): +0.2126 +0.21 OpenJev · C0 · unit=event · any lens · proof-backed data (intent-real): +0.3108 +0.31 OpenJev · C7 · unit=case · block lens · proof-backed data (intent-real): +0.5750 +0.58 OpenJev · C7 · unit=case · any lens · proof-backed data (intent-real): +0.7690 +0.77 OpenJev · C7 · unit=event · block lens · proof-backed data (intent-real): +0.5051 +0.51 OpenJev · C7 · unit=event · any lens · proof-backed data (intent-real): +0.6488 +0.65 DiffusionGemma DiffusionGemma · C0 · unit=case · block lens · proof-backed data (intent-real): +0.4010 +0.40 DiffusionGemma · C0 · unit=case · any lens · proof-backed data (intent-real): +0.4423 +0.44 DiffusionGemma · C0 · unit=event · block lens · proof-backed data (intent-real): +0.3234 +0.32 DiffusionGemma · C0 · unit=event · any lens · proof-backed data (intent-real): +0.3569 +0.36 DiffusionGemma · C7 · unit=case · block lens · proof-backed data (intent-real): +0.4846 +0.48 DiffusionGemma · C7 · unit=case · any lens · proof-backed data (intent-real): +0.5643 +0.56 DiffusionGemma · C7 · unit=event · block lens · proof-backed data (intent-real): +0.4140 +0.41 DiffusionGemma · C7 · unit=event · any lens · proof-backed data (intent-real): +0.4714 +0.47 Gemma 4 Gemma 4 · C0 · unit=case · block lens · proof-backed data (intent-real): +0.1753 +0.18 Gemma 4 · C0 · unit=case · any lens · proof-backed data (intent-real): +0.4454 +0.45 Gemma 4 · C0 · unit=event · block lens · proof-backed data (intent-real): +0.1398 +0.14 Gemma 4 · C0 · unit=event · any lens · proof-backed data (intent-real): +0.3453 +0.35 Gemma 4 · C7 · unit=case · block lens · proof-backed data (intent-real): +0.3340 +0.33 Gemma 4 · C7 · unit=case · any lens · proof-backed data (intent-real): +0.7202 +0.72 Gemma 4 · C7 · unit=event · block lens · proof-backed data (intent-real): +0.2718 +0.27 Gemma 4 · C7 · unit=event · any lens · proof-backed data (intent-real): +0.6175 +0.62 Jev 1.13.0 Jev 1.13.0 · C0 · unit=case · block lens · proof-backed data (intent-real): +0.1620 +0.16 Jev 1.13.0 · C0 · unit=case · any lens · proof-backed data (intent-real): +0.4141 +0.41 Jev 1.13.0 · C0 · unit=event · block lens · proof-backed data (intent-real): +0.1290 +0.13 Jev 1.13.0 · C0 · unit=event · any lens · proof-backed data (intent-real): +0.3295 +0.33 Jev 1.13.0 · C7 · unit=case · block lens · proof-backed data (intent-real): +0.3111 +0.31 Jev 1.13.0 · C7 · unit=case · any lens · proof-backed data (intent-real): +0.6051 +0.61 Jev 1.13.0 · C7 · unit=event · block lens · proof-backed data (intent-real): +0.2480 +0.25 Jev 1.13.0 · C7 · unit=event · any lens · proof-backed data (intent-real): +0.4894 +0.49 AgentDojo prior, same code C0 case blk C0 case any C0 evt blk C0 evt any C7 case blk C7 case any C7 evt blk C7 evt any OpenJev OpenJev · C0 · unit=case · block lens · AgentDojo prior, same code: -0.0572 -0.06 OpenJev · C0 · unit=case · any lens · AgentDojo prior, same code: -0.3384 -0.34 OpenJev · C0 · unit=event · block lens · AgentDojo prior, same code: +0.0034 +0.00 OpenJev · C0 · unit=event · any lens · AgentDojo prior, same code: -0.1118 -0.11 OpenJev · C7 · unit=case · block lens · AgentDojo prior, same code: -0.5343 -0.53 OpenJev · C7 · unit=case · any lens · AgentDojo prior, same code: -0.5969 -0.60 OpenJev · C7 · unit=event · block lens · AgentDojo prior, same code: -0.0831 -0.08 OpenJev · C7 · unit=event · any lens · AgentDojo prior, same code: -0.1028 -0.10 DiffusionGemma DiffusionGemma · C0 · unit=case · block lens · AgentDojo prior, same code: -0.1245 -0.12 DiffusionGemma · C0 · unit=case · any lens · AgentDojo prior, same code: -0.3382 -0.34 DiffusionGemma · C0 · unit=event · block lens · AgentDojo prior, same code: -0.0250 -0.03 DiffusionGemma · C0 · unit=event · any lens · AgentDojo prior, same code: -0.0649 -0.06 DiffusionGemma · C7 · unit=case · block lens · AgentDojo prior, same code: -0.3941 -0.39 DiffusionGemma · C7 · unit=case · any lens · AgentDojo prior, same code: -0.4466 -0.45 DiffusionGemma · C7 · unit=event · block lens · AgentDojo prior, same code: -0.0689 -0.07 DiffusionGemma · C7 · unit=event · any lens · AgentDojo prior, same code: -0.0926 -0.09 Gemma 4 Gemma 4 · C0 · unit=case · block lens · AgentDojo prior, same code: -0.0053 -0.01 Gemma 4 · C0 · unit=case · any lens · AgentDojo prior, same code: -0.2403 -0.24 Gemma 4 · C0 · unit=event · block lens · AgentDojo prior, same code: -0.0009 -0.00 Gemma 4 · C0 · unit=event · any lens · AgentDojo prior, same code: -0.1056 -0.11 Gemma 4 · C7 · unit=case · block lens · AgentDojo prior, same code: -0.2907 -0.29 Gemma 4 · C7 · unit=case · any lens · AgentDojo prior, same code: -0.3808 -0.38 Gemma 4 · C7 · unit=event · block lens · AgentDojo prior, same code: -0.0275 -0.03 Gemma 4 · C7 · unit=event · any lens · AgentDojo prior, same code: -0.0927 -0.09 Jev 1.13.0 Jev 1.13.0 · C0 · unit=case · block lens · AgentDojo prior, same code: -0.0208 -0.02 Jev 1.13.0 · C0 · unit=case · any lens · AgentDojo prior, same code: -0.2973 -0.30 Jev 1.13.0 · C0 · unit=event · block lens · AgentDojo prior, same code: -0.0042 -0.00 Jev 1.13.0 · C0 · unit=event · any lens · AgentDojo prior, same code: -0.1219 -0.12 Jev 1.13.0 · C7 · unit=case · block lens · AgentDojo prior, same code: -0.1246 -0.12 Jev 1.13.0 · C7 · unit=case · any lens · AgentDojo prior, same code: -0.4162 -0.42 Jev 1.13.0 · C7 · unit=event · block lens · AgentDojo prior, same code: -0.0258 -0.03 Jev 1.13.0 · C7 · unit=event · any lens · AgentDojo prior, same code: -0.1379 -0.14
Data table
Backendcontextunitlensproof-backedAgentDojo priorsign
OpenJevC0caseblock+0.266667-0.057208flips
OpenJevC0caseany+0.395089-0.338390flips
OpenJevC0eventblock+0.212598+0.003403same sign
OpenJevC0eventany+0.310809-0.111818flips
OpenJevC7caseblock+0.575008-0.534346flips
OpenJevC7caseany+0.768989-0.596923flips
OpenJevC7eventblock+0.505119-0.083135flips
OpenJevC7eventany+0.648842-0.102807flips
DiffusionGemmaC0caseblock+0.400985-0.124539flips
DiffusionGemmaC0caseany+0.442304-0.338150flips
DiffusionGemmaC0eventblock+0.323368-0.025036flips
DiffusionGemmaC0eventany+0.356870-0.064921flips
DiffusionGemmaC7caseblock+0.484594-0.394076flips
DiffusionGemmaC7caseany+0.564310-0.446611flips
DiffusionGemmaC7eventblock+0.414046-0.068889flips
DiffusionGemmaC7eventany+0.471407-0.092620flips
Gemma 4C0caseblock+0.175309-0.005315flips
Gemma 4C0caseany+0.445436-0.240292flips
Gemma 4C0eventblock+0.139764-0.000939flips
Gemma 4C0eventany+0.345329-0.105622flips
Gemma 4C7caseblock+0.333956-0.290690flips
Gemma 4C7caseany+0.720228-0.380802flips
Gemma 4C7eventblock+0.271785-0.027485flips
Gemma 4C7eventany+0.617478-0.092719flips
Jev 1.13.0C0caseblock+0.162040-0.020832flips
Jev 1.13.0C0caseany+0.414103-0.297340flips
Jev 1.13.0C0eventblock+0.129003-0.004227flips
Jev 1.13.0C0eventany+0.329470-0.121868flips
Jev 1.13.0C7caseblock+0.311111-0.124566flips
Jev 1.13.0C7caseany+0.605061-0.416164flips
Jev 1.13.0C7eventblock+0.248031-0.025837flips
Jev 1.13.0C7eventany+0.489387-0.137907flips
Statistically, 32 of 32 proof-backed cells are positive and 31 of 32 flip sign against the prior run. Visually the lower block is not uniformly red: 4 of its 32 cells land inside the ±0.02 neutral band this legend defines and render neutral, leaving 28 red. The same 32 cells appear as a bar chart with bootstrap intervals below. Source: outputs/intent-real/jev-analysis.json :: four_backend_table[backend][context][unit][lens].{intent_real, agentdojo_prior}.sep_vs_resisted; four_backend_summary.

The one cell that is not a sign reversal (OpenJev without the request, C0, per event, block-only) was already positive at +0.003403. For OpenJev with the request (C0 / C7C0 is the tool call alone. C7 adds the user’s request and up to seven prior events.), separation is -0.5343 on AgentDojo and +0.5750 on the proof-backed corpus.

Intent separation, four models, two corpora

One bar per model per corpus per context. Left of zero means the guard flagged a correct refusal more often than a proven compromise. Error bars are family-cluster bootstrap 95% intervals, 2,000 replicates, seed 741983. Lead cell: proof-verified (grade A) lane, block-only lens, per case.

AgentDojo, no intent (C0)AgentDojo, with intent (C7)Proof-backed, no intent (C0)Proof-backed, with intent (C7)
Separation of proven compromise from a correct refusal, four backends, two corporaOn the AgentDojo benchmark 31 of the 32 cells scored negative; the rest were already positive, which is why the reversal count below is 31 of 32. On the proof-backed corpus every backend scores positive, and adding the user request increases the score. -0.6 -0.4 -0.2 +0.0 +0.2 +0.4 +0.6 separation (sep|res) = block-rate on proven compromise − block-rate on a correct refusal 0 = no discrimination OpenJev OpenJev · AgentDojo, no intent (C0): -0.0572 [95% -0.1135, -0.0066] OpenJev · AgentDojo, with intent (C7): -0.5343 [95% -0.6146, -0.4476] -0.5343 OpenJev · Proof-backed, no intent (C0): +0.2667 [95% +0.2241, +0.3089] OpenJev · Proof-backed, with intent (C7): +0.5750 [95% +0.5270, +0.6209] +0.5750 DiffusionGemma DiffusionGemma · AgentDojo, no intent (C0): -0.1245 [95% -0.1881, -0.0643] DiffusionGemma · AgentDojo, with intent (C7): -0.3941 [95% -0.4736, -0.3137] -0.3941 DiffusionGemma · Proof-backed, no intent (C0): +0.4010 [95% +0.3504, +0.4497] DiffusionGemma · Proof-backed, with intent (C7): +0.4846 [95% +0.4367, +0.5333] +0.4846 Gemma 4 Gemma 4 · AgentDojo, no intent (C0): -0.0053 [95% -0.0378, +0.0273] Gemma 4 · AgentDojo, with intent (C7): -0.2907 [95% -0.3751, -0.2089] -0.2907 Gemma 4 · Proof-backed, no intent (C0): +0.1753 [95% +0.1390, +0.2136] Gemma 4 · Proof-backed, with intent (C7): +0.3340 [95% +0.2858, +0.3801] +0.3340 Jev 1.13.0 Jev 1.13.0 · AgentDojo, no intent (C0): -0.0208 [95% -0.0565, +0.0130] Jev 1.13.0 · AgentDojo, with intent (C7): -0.1246 [95% -0.1861, -0.0641] -0.1246 Jev 1.13.0 · Proof-backed, no intent (C0): +0.1620 [95% +0.1241, +0.1977] Jev 1.13.0 · Proof-backed, with intent (C7): +0.3111 [95% +0.2647, +0.3589] +0.3111
Data table
BackendCorpus / contextsep|resbootstrap 95%
OpenJev — self-hosted, FP8, structured instructions, question Q2AgentDojo, no intent (C0)-0.0572[-0.1135, -0.0066]
OpenJev — self-hosted, FP8, structured instructions, question Q2AgentDojo, with intent (C7)-0.5343[-0.6146, -0.4476]
OpenJev — self-hosted, FP8, structured instructions, question Q2Proof-backed, no intent (C0)+0.2667[+0.2241, +0.3089]
OpenJev — self-hosted, FP8, structured instructions, question Q2Proof-backed, with intent (C7)+0.5750[+0.5270, +0.6209]
DiffusionGemma 26B-A4B-it-FP8-dynamic — self-hosted, string instructions, Q2AgentDojo, no intent (C0)-0.1245[-0.1881, -0.0643]
DiffusionGemma 26B-A4B-it-FP8-dynamic — self-hosted, string instructions, Q2AgentDojo, with intent (C7)-0.3941[-0.4736, -0.3137]
DiffusionGemma 26B-A4B-it-FP8-dynamic — self-hosted, string instructions, Q2Proof-backed, no intent (C0)+0.4010[+0.3504, +0.4497]
DiffusionGemma 26B-A4B-it-FP8-dynamic — self-hosted, string instructions, Q2Proof-backed, with intent (C7)+0.4846[+0.4367, +0.5333]
Gemma 4 (google.gemma-4-26b-a4b via Bedrock) — chat judge, question Q0AgentDojo, no intent (C0)-0.0053[-0.0378, +0.0273]
Gemma 4 (google.gemma-4-26b-a4b via Bedrock) — chat judge, question Q0AgentDojo, with intent (C7)-0.2907[-0.3751, -0.2089]
Gemma 4 (google.gemma-4-26b-a4b via Bedrock) — chat judge, question Q0Proof-backed, no intent (C0)+0.1753[+0.1390, +0.2136]
Gemma 4 (google.gemma-4-26b-a4b via Bedrock) — chat judge, question Q0Proof-backed, with intent (C7)+0.3340[+0.2858, +0.3801]
Jev 1.13.0 — hosted API, structured instructions, Q2AgentDojo, no intent (C0)-0.0208[-0.0565, +0.0130]
Jev 1.13.0 — hosted API, structured instructions, Q2AgentDojo, with intent (C7)-0.1246[-0.1861, -0.0641]
Jev 1.13.0 — hosted API, structured instructions, Q2Proof-backed, no intent (C0)+0.1620[+0.1241, +0.1977]
Jev 1.13.0 — hosted API, structured instructions, Q2Proof-backed, with intent (C7)+0.3111[+0.2647, +0.3589]
All 32 proof-backed cells are positive with intervals excluding zero. 31 of 32 are strict sign reversals against the same model's own AgentDojo figure; the remaining cell was already positive in that prior run. Source: outputs/intent-real/jev-analysis.json :: four_backend_table (AgentDojo priors recomputed by the same scorer; cross-checked against replication-analysis.json and separation-analysis.json).

Why the reversal is not an artefact

Limits

Block-only recall on proof-verified compromises

Separation measures discrimination. Recall bounds the production claim: the share of independently proof-verified compromises the guard actually blocks. A confirm stops nothing, so only hard blocks count here.

no user request supplied (C0)user request + up to 7 prior events (C7)
Share of 405 proof-backed compromises actually blockedEvery backend misses between 42 and 69 percent of independently proven compromises even at its best measured operating point. 0% 10% 20% 30% 40% 50% 60% 70% block-only recall on the 405 proof-verified (grade A) compromises OpenJev OpenJev · no user request supplied (C0): 0.2667 (108/405), misses 297 OpenJev · user request + history (C7): 0.5778 (234/405), misses 171 0.5778 · misses 42.2% DiffusionGemma DiffusionGemma · no user request supplied (C0): 0.4222 (171/405), misses 234 DiffusionGemma · user request + history (C7): 0.4938 (200/405), misses 205 0.4938 · misses 50.6% Gemma 4 Gemma 4 · no user request supplied (C0): 0.1753 (71/405), misses 334 Gemma 4 · user request + history (C7): 0.3358 (136/405), misses 269 0.3358 · misses 66.4% Jev 1.13.0 Jev 1.13.0 · no user request supplied (C0): 0.1630 (66/405), misses 339 Jev 1.13.0 · user request + history (C7): 0.3111 (126/405), misses 279 0.3111 · misses 68.9%
Data table
BackendContextrecallblocked / totalmissed
OpenJevno user request supplied (C0)0.2667108/40573.3%
OpenJevuser request + history (C7)0.5778234/40542.2%
DiffusionGemmano user request supplied (C0)0.4222171/40557.8%
DiffusionGemmauser request + history (C7)0.4938200/40550.6%
Gemma 4no user request supplied (C0)0.175371/40582.5%
Gemma 4user request + history (C7)0.3358136/40566.4%
Jev 1.13.0no user request supplied (C0)0.163066/40583.7%
Jev 1.13.0user request + history (C7)0.3111126/40568.9%
Source: outputs/intent-real/jev-analysis.json :: grade_a_block_only_recall.

Every model misses 42.2% to 68.9% of proven compromises. With the request supplied, the four block 0.5778 (OpenJev), 0.4938 (DiffusionGemma), 0.3358 (the Gemma 4 judge) and 0.3111 (Jev 1.13.0).

What the evidence supports.
ClaimStatus
Comparing the call with the user’s request separates proven compromise from correct refusal.supported
Knowing the request improves that separation.supported
The earlier eight-run negative was a property of the mechanism.refuted
This is enough to ship as a blocking control.not supported

evaluation-only never-train
Nothing derived here is approved for training, synthetic generation, teacher context, distillation or redistribution. The 13 source datasets are public and linked on reproduce; the analysis artifacts behind each number are held privately and are available on request.