Which small model to put in front of the judge
A tool-call guard here is a cascade: deterministic rules, then a small model, then the Gemma 4
judge on the cases the small model is unsure of. Each small model answers allow,
confirm or block for one agent tool call. This page ranks the
candidates for the middle tier on the Broad comparison corpus: 3,817 scorable
cases, 436 of them unsafe and 3,381 benign. Every model is held to one
shared budgetOne block false-positive rate every model is held to. Each model’s threshold is set so it makes no more false blocks than OpenJev makes on this corpus as shipped, then its recall and F1 are read at that point..
The answer
rules → OpenJev → Gemma 4 judge, two-sided routingThe small model settles a case alone when it is confident either way; only the uncertain band between the allow and block thresholds goes to the judge. at an allow threshold of 0.30, with escalate-on-confirmWhen the rule engine answers confirm, that answer becomes a floor and the later tiers still run, so a case the small model or the judge would block ends as a block.: block-only F1 0.75173 at block FPR 0.00414 (14 false blocks), with the judge called on 15.95% of cases. The judge alone scores 0.71248 at block FPR 0.00680. OpenJev weights are CC BY-NC 4.0: non-commercial, attribution required, contact the authors for commercial use.allow ends the cascade. This is how the cascade runs today; an advisory rule confirm therefore caps a case a later tier would block. on any rule-engine answer, scores 0.73773. The two differ only on the 8 cases where an advisory rule confirm ended the cascade before a later tier could block: 263 true blocks become 271, at the same 14 false blocks and the same judge-call rate. The recommendation is the escalating version; the cost panels on Deploy are measured on the version that runs today.rules → OpenJev with no judge scores 0.21285 against 0.08215 for the full cascade and 0.10980 for the judge alone. See Deploy.Every model at the shared budget
One ranking. Each model’s threshold is set so it makes no more than 13
false blocks, and its block-only F1F1 counting only a hard block as a catch. A confirm counts as a miss, because a confirm does not stop the tool call., recall and precision are read there. The shipped
column is the same model at the threshold it ships with, which mixes the model with its
calibration; it is printed for reference and does not rank anything.
| Rank | Model | License | F1 at the budget 95% interval | recall | precision | true / false blocks | AUC | block-only F1 as shipped | parameters |
|---|---|---|---|---|---|---|---|---|---|
| 1 | open-jev-qwen-27b | apache-2.0 | 0.73876 0.70058–0.77309 | 0.60321 | 0.95290 | 263 / 13 | 0.94803 | 0.33206 | 27B (nominal) |
| 2 | OpenJev | CC BY-NC 4.0 (non-commercial) | 0.71162 0.67169–0.75034 | 0.56881 | 0.95019 | 248 / 13 | 0.97454 | 0.70231 | not recorded |
| 3 | Jev 1.13.0 | commercial API, $0.042/M input tokens | 0.58360 0.53443–0.62829 | 0.42431 | 0.93434 | 185 / 13 | 0.93942 | 0.54153 | not recorded |
| 4 | gemma-4-26B-A4B-it | apache-2.0 | 0.32463 0.27206–0.37113 | 0.19954 | 0.87000 | 87 / 13 | 0.88550 | 0.47496 | 25.8B (counted) |
| 5 | jevify-gemma4-26b-a4b | gemma | 0.29278 0.24231–0.34019 | 0.17661 | 0.85556 | 77 / 13 | 0.87989 | 0.04922 | 25.8B (counted) |
| 6 | open-jev-qwen-9b | apache-2.0 | 0.24609 0.19706–0.29769 | 0.14450 | 0.82895 | 63 / 13 | 0.78269 | 0.17551 | 9B (nominal) |
| 7 | DiffusionGemma | apache-2.0 | 0.22530 0.17923–0.27495 | 0.13073 | 0.81429 | 57 / 13 | 0.77729 | 0.26792 | 26B (nominal) |
| 8 | kev-9b | apache-2.0 | 0.19316 0.14729–0.23886 | 0.11009 | 0.78689 | 48 / 13 | 0.82622 | 0.01814 | 9B (nominal) |
| 9 | open-jev-qwen-2b | apache-2.0 | 0.08120 0.04785–0.11861 | 0.04358 | 0.59375 | 19 / 13 | 0.55352 | 0.17195 | 2B (nominal) |
| 10 | decider-2b re-mining only | not recorded | 0.07296 0.04184–0.10856 | 0.03899 | 0.56667 | 17 / 13 | 0.77040 | 0.10843 | 2B (nominal) |
| 11 | bespoke-nimble-9b | apache-2.0 | 0.00887 0.00000–0.02278 | 0.00459 | 0.13333 | 2 / 13 | 0.89630 | 0.21569 | 9B (nominal) |
| 12 | SecJudge | apache-2.0 | 0.00000 0.00000–0.00000 | 0.00000 | 0.00000 | 0 / 1 | 0.65324 | 0.20724 | 396M (counted) |
| ref | Gemma 4 judge reference | apache-2.0 weights, served via Bedrock (paid service) | — | — | — | — | — | 0.71248 | not recorded |
Every ranked model answered the same question format, Q2The question format every ranked model answered: one disposition choice (allow / confirm / block) plus two scores., except SecJudge, a
classifier that takes one serialised string and answers no question at all. Question format
moves F1 about as much as the choice of model does, which is why the ranking holds it fixed;
the measurement is on Method.
Precision, recall and F1 per model at the shared budget
Data table
| Model | precision | precision, Wilson 95% | recall | recall, Wilson 95% | F1 | F1, bootstrap 95% |
|---|---|---|---|---|---|---|
| open-jev-qwen-27b | 0.95290 | 0.92109 to 0.97227 | 0.60321 | 0.55658 to 0.64804 | 0.73876 | 0.70058 to 0.77309 |
| OpenJev | 0.95019 | 0.91666 to 0.97066 | 0.56881 | 0.52192 to 0.61449 | 0.71162 | 0.67169 to 0.75034 |
| Jev 1.13.0 | 0.93434 | 0.89092 to 0.96123 | 0.42431 | 0.37878 to 0.47117 | 0.58360 | 0.53443 to 0.62829 |
| gemma-4-26B-A4B-it | 0.87000 | 0.79020 to 0.92243 | 0.19954 | 0.16472 to 0.23961 | 0.32463 | 0.27206 to 0.37113 |
| jevify-gemma4-26b-a4b | 0.85556 | 0.76840 to 0.91360 | 0.17661 | 0.14368 to 0.21518 | 0.29278 | 0.24231 to 0.34019 |
| open-jev-qwen-9b | 0.82895 | 0.72902 to 0.89722 | 0.14450 | 0.11460 to 0.18060 | 0.24609 | 0.19706 to 0.29769 |
| DiffusionGemma | 0.81429 | 0.70774 to 0.88813 | 0.13073 | 0.10229 to 0.16563 | 0.22530 | 0.17923 to 0.27495 |
| kev-9b | 0.78689 | 0.66878 to 0.87100 | 0.11009 | 0.08405 to 0.14295 | 0.19316 | 0.14729 to 0.23886 |
| open-jev-qwen-2b | 0.59375 | 0.42260 to 0.74480 | 0.04358 | 0.02807 to 0.06706 | 0.08120 | 0.04785 to 0.11861 |
| decider-2b | 0.56667 | 0.39197 to 0.72623 | 0.03899 | 0.02448 to 0.06155 | 0.07296 | 0.04184 to 0.10856 |
| bespoke-nimble-9b | 0.13333 | 0.03736 to 0.37882 | 0.00458716 | 0.00125887 to 0.01657 | 0.00886918 | 0.00000000 to 0.02278 |
| SecJudge | 0.00000000 | 0.00000000 to 0.79345 | 0.00000000 | 0.00000000 to 0.00873374 | 0.00000000 | 0.00000000 to 0.00000000 |
outputs/curves/curves-s2.json :: arms.<arm>.at_budget.Every model at every other budget
The table is one point on each model’s precision-recall curve. A deployment that tolerates more false blocks reads a different point, and the order can change.
Precision against recall per model, with prevalence drawn
Prevalence 0.11423 is the horizontal baseline on every panel.
Data table
| Model | points plotted | precision | recall | F1 | precision, Wilson 95% | recall, Wilson 95% | F1, bootstrap 95% |
|---|---|---|---|---|---|---|---|
| open-jev-qwen-27b | 411 | 0.95290 | 0.60321 | 0.73876 | 0.92109 to 0.97227 | 0.55658 to 0.64804 | 0.70058 to 0.77309 |
| OpenJev | 303 | 0.95019 | 0.56881 | 0.71162 | 0.91666 to 0.97066 | 0.52192 to 0.61449 | 0.67169 to 0.75034 |
| Jev 1.13.0 | 93 | 0.93434 | 0.42431 | 0.58360 | 0.89092 to 0.96123 | 0.37878 to 0.47117 | 0.53443 to 0.62829 |
| gemma-4-26B-A4B-it | 265 | 0.87000 | 0.19954 | 0.32463 | 0.79020 to 0.92243 | 0.16472 to 0.23961 | 0.27206 to 0.37113 |
| jevify-gemma4-26b-a4b | 279 | 0.85556 | 0.17661 | 0.29278 | 0.76840 to 0.91360 | 0.14368 to 0.21518 | 0.24231 to 0.34019 |
| open-jev-qwen-9b | 205 | 0.82895 | 0.14450 | 0.24609 | 0.72902 to 0.89722 | 0.11460 to 0.18060 | 0.19706 to 0.29769 |
| DiffusionGemma | 238 | 0.81429 | 0.13073 | 0.22530 | 0.70774 to 0.88813 | 0.10229 to 0.16563 | 0.17923 to 0.27495 |
| kev-9b | 205 | 0.78689 | 0.11009 | 0.19316 | 0.66878 to 0.87100 | 0.08405 to 0.14295 | 0.14729 to 0.23886 |
| open-jev-qwen-2b | 152 | 0.59375 | 0.04358 | 0.08120 | 0.42260 to 0.74480 | 0.02807 to 0.06706 | 0.04785 to 0.11861 |
| decider-2b | 163 | 0.56667 | 0.03899 | 0.07296 | 0.39197 to 0.72623 | 0.02448 to 0.06155 | 0.04184 to 0.10856 |
| bespoke-nimble-9b | 145 | 0.13333 | 0.00458716 | 0.00886918 | 0.03736 to 0.37882 | 0.00125887 to 0.01657 | 0.00000000 to 0.02278 |
| SecJudge | 98 | 0.00000000 | 0.00000000 | 0.00000000 | 0.00000000 to 0.79345 | 0.00000000 to 0.00873374 | 0.00000000 to 0.00000000 |
strata.split_group is unique per case, so its 3,817 families are its 3,817 scorable cases and the family bootstrap is a case bootstrap. Source: outputs/curves/curves-s2.json :: arms.<arm>.pr, arms.<arm>.at_budget and corpus.prevalence.Size against score
Parameter counts are counted from the served weights where the artifacts record them and nominal (read from the base model’s name) otherwise; 2 (OpenJev and Jev 1.13.0) carry neither and are not plotted.
Parameters against F1 at the shared budget
3 counted figures and 7 nominal ones.
Data table
| Model | parameters | basis | band | F1 at the shared budget |
|---|---|---|---|---|
| SecJudge | 395,836,421 | counted | under 3B | 0.00000000 |
| open-jev-qwen-2b | 2,000,000,000 | nominal | under 3B | 0.08120 |
| decider-2b | 2,000,000,000 | nominal | under 3B | 0.07296 |
| open-jev-qwen-9b | 9,000,000,000 | nominal | 6B and up | 0.24609 |
| kev-9b | 9,000,000,000 | nominal | 6B and up | 0.19316 |
| bespoke-nimble-9b | 9,000,000,000 | nominal | 6B and up | 0.00886918 |
| gemma-4-26B-A4B-it | 25,805,936,206 | counted | 6B and up | 0.32463 |
| jevify-gemma4-26b-a4b | 25,805,936,206 | counted | 6B and up | 0.29278 |
| DiffusionGemma | 26,000,000,000 | nominal | 6B and up | 0.22530 |
| open-jev-qwen-27b | 27,000,000,000 | nominal | 6B and up | 0.73876 |
each arm's serving record, outputs/gemma4jev/artifacts/out/jevify-weight-diff.json, and outputs/curves/curves-s2.json :: arms.<arm>.at_budget.f1; except the log position of each point, and the hollow fill that marks a nominal figure.Does the score mean what it says
A threshold is only portable if the probability behind it is calibrated. Each panel buckets the per-case score and plots the share of the bucket that is unsafe, with the case count beside it.
Predicted probability against observed positive rate, per model
10 equal-width buckets of the per-case maximum P(block), with each bucket’s case count printed.
Data table
| Model | bucket | cases | positives | mean predicted | observed rate | observed rate, Wilson 95% |
|---|---|---|---|---|---|---|
| open-jev-qwen-27b | 0.0 to 0.1 | 3,390 | 84 | 0.01977 | 0.02478 | 0.02006 to 0.03057 |
| open-jev-qwen-27b | 0.1 to 0.2 | 166 | 102 | 0.14237 | 0.61446 | 0.53862 to 0.68511 |
| open-jev-qwen-27b | 0.2 to 0.3 | 109 | 100 | 0.24438 | 0.91743 | 0.85049 to 0.95595 |
| open-jev-qwen-27b | 0.3 to 0.4 | 85 | 84 | 0.35248 | 0.98824 | 0.93633 to 0.99792 |
| open-jev-qwen-27b | 0.4 to 0.5 | 30 | 30 | 0.44028 | 1 | 0.88649 to 1 |
| open-jev-qwen-27b | 0.5 to 0.6 | 22 | 21 | 0.54063 | 0.95455 | 0.78202 to 0.99193 |
| open-jev-qwen-27b | 0.6 to 0.7 | 4 | 4 | 0.64929 | 1 | 0.51011 to 1 |
| open-jev-qwen-27b | 0.7 to 0.8 | 3 | 3 | 0.77815 | 1 | 0.43850 to 1 |
| open-jev-qwen-27b | 0.8 to 0.9 | 6 | 6 | 0.86237 | 1 | 0.60967 to 1 |
| open-jev-qwen-27b | 0.9 to 1.0 | 2 | 2 | 0.94717 | 1 | 0.34238 to 1 |
| OpenJev | 0.0 to 0.1 | 3,372 | 81 | 0.01313 | 0.02402 | 0.01937 to 0.02976 |
| OpenJev | 0.1 to 0.2 | 88 | 48 | 0.14455 | 0.54545 | 0.44170 to 0.64541 |
| OpenJev | 0.2 to 0.3 | 64 | 32 | 0.25181 | 0.50000 | 0.38102 to 0.61898 |
| OpenJev | 0.3 to 0.4 | 39 | 33 | 0.34620 | 0.84615 | 0.70271 to 0.92753 |
| OpenJev | 0.4 to 0.5 | 49 | 43 | 0.45022 | 0.87755 | 0.75756 to 0.94265 |
| OpenJev | 0.5 to 0.6 | 42 | 38 | 0.54498 | 0.90476 | 0.77935 to 0.96234 |
| OpenJev | 0.6 to 0.7 | 47 | 46 | 0.65070 | 0.97872 | 0.88887 to 0.99623 |
| OpenJev | 0.7 to 0.8 | 34 | 34 | 0.75730 | 1 | 0.89849 to 1 |
| OpenJev | 0.8 to 0.9 | 58 | 57 | 0.84549 | 0.98276 | 0.90859 to 0.99695 |
| OpenJev | 0.9 to 1.0 | 24 | 24 | 0.94472 | 1 | 0.86202 to 1 |
| Jev 1.13.0 | 0.0 to 0.1 | 3,426 | 102 | 0.00514302 | 0.02977 | 0.02459 to 0.03601 |
| Jev 1.13.0 | 0.1 to 0.2 | 87 | 61 | 0.14517 | 0.70115 | 0.59813 to 0.78716 |
| Jev 1.13.0 | 0.2 to 0.3 | 73 | 59 | 0.24096 | 0.80822 | 0.70344 to 0.88218 |
| Jev 1.13.0 | 0.3 to 0.4 | 68 | 59 | 0.34588 | 0.86765 | 0.76720 to 0.92878 |
| Jev 1.13.0 | 0.4 to 0.5 | 58 | 52 | 0.44121 | 0.89655 | 0.79212 to 0.95172 |
| Jev 1.13.0 | 0.5 to 0.6 | 41 | 41 | 0.54415 | 1 | 0.91433 to 1 |
| Jev 1.13.0 | 0.6 to 0.7 | 30 | 29 | 0.64133 | 0.96667 | 0.83330 to 0.99409 |
| Jev 1.13.0 | 0.7 to 0.8 | 16 | 15 | 0.73188 | 0.93750 | 0.71671 to 0.98888 |
| Jev 1.13.0 | 0.8 to 0.9 | 8 | 8 | 0.84125 | 1 | 0.67559 to 1 |
| Jev 1.13.0 | 0.9 to 1.0 | 10 | 10 | 0.95000 | 1 | 0.72247 to 1 |
| gemma-4-26B-A4B-it | 0.0 to 0.1 | 3,003 | 112 | 0.05352 | 0.03730 | 0.03109 to 0.04469 |
| gemma-4-26B-A4B-it | 0.1 to 0.2 | 529 | 125 | 0.12817 | 0.23629 | 0.20208 to 0.27432 |
| gemma-4-26B-A4B-it | 0.2 to 0.3 | 65 | 34 | 0.24006 | 0.52308 | 0.40380 to 0.63978 |
| gemma-4-26B-A4B-it | 0.3 to 0.4 | 31 | 16 | 0.34656 | 0.51613 | 0.34840 to 0.68030 |
| gemma-4-26B-A4B-it | 0.4 to 0.5 | 38 | 20 | 0.44194 | 0.52632 | 0.37259 to 0.67521 |
| gemma-4-26B-A4B-it | 0.5 to 0.6 | 31 | 26 | 0.54915 | 0.83871 | 0.67366 to 0.92907 |
| gemma-4-26B-A4B-it | 0.6 to 0.7 | 42 | 35 | 0.65084 | 0.83333 | 0.69396 to 0.91684 |
| gemma-4-26B-A4B-it | 0.7 to 0.8 | 35 | 29 | 0.74933 | 0.82857 | 0.67318 to 0.91897 |
| gemma-4-26B-A4B-it | 0.8 to 0.9 | 35 | 31 | 0.83584 | 0.88571 | 0.74049 to 0.95465 |
| gemma-4-26B-A4B-it | 0.9 to 1.0 | 8 | 8 | 0.92490 | 1 | 0.67559 to 1 |
| jevify-gemma4-26b-a4b | 0.0 to 0.1 | 3,773 | 396 | 0.00292870 | 0.10496 | 0.09557 to 0.11514 |
| jevify-gemma4-26b-a4b | 0.1 to 0.2 | 20 | 17 | 0.12455 | 0.85000 | 0.63958 to 0.94763 |
| jevify-gemma4-26b-a4b | 0.2 to 0.3 | 10 | 10 | 0.24279 | 1 | 0.72247 to 1 |
| jevify-gemma4-26b-a4b | 0.3 to 0.4 | 3 | 2 | 0.32102 | 0.66667 | 0.20766 to 0.93851 |
| jevify-gemma4-26b-a4b | 0.4 to 0.5 | 5 | 5 | 0.43032 | 1 | 0.56552 to 1 |
| jevify-gemma4-26b-a4b | 0.5 to 0.6 | 1 | 1 | 0.56031 | 1 | 0.20655 to 1 |
| jevify-gemma4-26b-a4b | 0.6 to 0.7 | 1 | 1 | 0.67752 | 1 | 0.20655 to 1 |
| jevify-gemma4-26b-a4b | 0.7 to 0.8 | 1 | 1 | 0.79245 | 1 | 0.20655 to 1 |
| jevify-gemma4-26b-a4b | 0.9 to 1.0 | 3 | 3 | 0.95399 | 1 | 0.43850 to 1 |
| open-jev-qwen-9b | 0.0 to 0.1 | 3,679 | 341 | 0.02260 | 0.09269 | 0.08374 to 0.10249 |
| open-jev-qwen-9b | 0.1 to 0.2 | 59 | 32 | 0.13512 | 0.54237 | 0.41658 to 0.66299 |
| open-jev-qwen-9b | 0.2 to 0.3 | 17 | 13 | 0.24563 | 0.76471 | 0.52738 to 0.90445 |
| open-jev-qwen-9b | 0.3 to 0.4 | 10 | 8 | 0.33785 | 0.80000 | 0.49016 to 0.94332 |
| open-jev-qwen-9b | 0.4 to 0.5 | 8 | 4 | 0.44196 | 0.50000 | 0.21522 to 0.78478 |
| open-jev-qwen-9b | 0.5 to 0.6 | 8 | 7 | 0.54878 | 0.87500 | 0.52911 to 0.97758 |
| open-jev-qwen-9b | 0.6 to 0.7 | 3 | 2 | 0.65367 | 0.66667 | 0.20766 to 0.93851 |
| open-jev-qwen-9b | 0.7 to 0.8 | 11 | 10 | 0.77306 | 0.90909 | 0.62264 to 0.98377 |
| open-jev-qwen-9b | 0.8 to 0.9 | 5 | 5 | 0.84864 | 1 | 0.56552 to 1 |
| open-jev-qwen-9b | 0.9 to 1.0 | 17 | 14 | 0.95487 | 0.82353 | 0.58971 to 0.93809 |
| DiffusionGemma | 0.0 to 0.1 | 3,547 | 263 | 0.00826698 | 0.07415 | 0.06598 to 0.08324 |
| DiffusionGemma | 0.1 to 0.2 | 81 | 42 | 0.14096 | 0.51852 | 0.41136 to 0.62400 |
| DiffusionGemma | 0.2 to 0.3 | 52 | 29 | 0.25021 | 0.55769 | 0.42340 to 0.68405 |
| DiffusionGemma | 0.3 to 0.4 | 40 | 28 | 0.34126 | 0.70000 | 0.54570 to 0.81925 |
| DiffusionGemma | 0.4 to 0.5 | 21 | 15 | 0.45379 | 0.71429 | 0.50044 to 0.86186 |
| DiffusionGemma | 0.5 to 0.6 | 18 | 13 | 0.53993 | 0.72222 | 0.49127 to 0.87500 |
| DiffusionGemma | 0.6 to 0.7 | 13 | 8 | 0.66025 | 0.61538 | 0.35523 to 0.82290 |
| DiffusionGemma | 0.7 to 0.8 | 24 | 18 | 0.75377 | 0.75000 | 0.55101 to 0.88001 |
| DiffusionGemma | 0.8 to 0.9 | 8 | 7 | 0.83024 | 0.87500 | 0.52911 to 0.97758 |
| DiffusionGemma | 0.9 to 1.0 | 13 | 13 | 0.96392 | 1 | 0.77190 to 1 |
| kev-9b | 0.0 to 0.1 | 3,392 | 264 | 0.04014 | 0.07783 | 0.06928 to 0.08733 |
| kev-9b | 0.1 to 0.2 | 371 | 130 | 0.12893 | 0.35040 | 0.30361 to 0.40026 |
| kev-9b | 0.2 to 0.3 | 44 | 36 | 0.22838 | 0.81818 | 0.68039 to 0.90487 |
| kev-9b | 0.3 to 0.4 | 4 | 1 | 0.33643 | 0.25000 | 0.04559 to 0.69936 |
| kev-9b | 0.4 to 0.5 | 1 | 1 | 0.40080 | 1 | 0.20655 to 1 |
| kev-9b | 0.5 to 0.6 | 2 | 1 | 0.57687 | 0.50000 | 0.09453 to 0.90547 |
| kev-9b | 0.6 to 0.7 | 2 | 2 | 0.66204 | 1 | 0.34238 to 1 |
| kev-9b | 0.9 to 1.0 | 1 | 1 | 0.95163 | 1 | 0.20655 to 1 |
| open-jev-qwen-2b | 0.0 to 0.1 | 3,273 | 337 | 0.03709 | 0.10296 | 0.09301 to 0.11385 |
| open-jev-qwen-2b | 0.1 to 0.2 | 208 | 31 | 0.13822 | 0.14904 | 0.10703 to 0.20378 |
| open-jev-qwen-2b | 0.2 to 0.3 | 59 | 7 | 0.24581 | 0.11864 | 0.05868 to 0.22524 |
| open-jev-qwen-2b | 0.3 to 0.4 | 40 | 3 | 0.35053 | 0.07500 | 0.02584 to 0.19864 |
| open-jev-qwen-2b | 0.4 to 0.5 | 40 | 3 | 0.44373 | 0.07500 | 0.02584 to 0.19864 |
| open-jev-qwen-2b | 0.5 to 0.6 | 47 | 7 | 0.54262 | 0.14894 | 0.07407 to 0.27686 |
| open-jev-qwen-2b | 0.6 to 0.7 | 45 | 8 | 0.65060 | 0.17778 | 0.09294 to 0.31330 |
| open-jev-qwen-2b | 0.7 to 0.8 | 35 | 8 | 0.75057 | 0.22857 | 0.12066 to 0.39017 |
| open-jev-qwen-2b | 0.8 to 0.9 | 28 | 10 | 0.85511 | 0.35714 | 0.20706 to 0.54170 |
| open-jev-qwen-2b | 0.9 to 1.0 | 42 | 22 | 0.94761 | 0.52381 | 0.37722 to 0.66640 |
| decider-2b | 0.0 to 0.1 | 3,267 | 251 | 0.04746 | 0.07683 | 0.06819 to 0.08647 |
| decider-2b | 0.1 to 0.2 | 379 | 111 | 0.13346 | 0.29288 | 0.24932 to 0.34059 |
| decider-2b | 0.2 to 0.3 | 72 | 27 | 0.24231 | 0.37500 | 0.27219 to 0.49047 |
| decider-2b | 0.3 to 0.4 | 30 | 16 | 0.34378 | 0.53333 | 0.36142 to 0.69768 |
| decider-2b | 0.4 to 0.5 | 13 | 6 | 0.44099 | 0.46154 | 0.23206 to 0.70856 |
| decider-2b | 0.5 to 0.6 | 18 | 6 | 0.55396 | 0.33333 | 0.16279 to 0.56251 |
| decider-2b | 0.6 to 0.7 | 16 | 8 | 0.65500 | 0.50000 | 0.28000 to 0.72000 |
| decider-2b | 0.7 to 0.8 | 11 | 5 | 0.73179 | 0.45455 | 0.21271 to 0.71991 |
| decider-2b | 0.8 to 0.9 | 11 | 6 | 0.85790 | 0.54545 | 0.28009 to 0.78729 |
| bespoke-nimble-9b | 0.0 to 0.1 | 9 | 0 | 0.08758 | 0.00000000 | 0.00000000 to 0.29915 |
| bespoke-nimble-9b | 0.1 to 0.2 | 2,237 | 21 | 0.15781 | 0.00938757 | 0.00614827 to 0.01431 |
| bespoke-nimble-9b | 0.2 to 0.3 | 1,074 | 150 | 0.23817 | 0.13966 | 0.12022 to 0.16168 |
| bespoke-nimble-9b | 0.3 to 0.4 | 325 | 183 | 0.34640 | 0.56308 | 0.50873 to 0.61595 |
| bespoke-nimble-9b | 0.4 to 0.5 | 118 | 68 | 0.44106 | 0.57627 | 0.48609 to 0.66164 |
| bespoke-nimble-9b | 0.5 to 0.6 | 38 | 12 | 0.53819 | 0.31579 | 0.19085 to 0.47456 |
| bespoke-nimble-9b | 0.6 to 0.7 | 12 | 2 | 0.63814 | 0.16667 | 0.04697 to 0.44803 |
| bespoke-nimble-9b | 0.7 to 0.8 | 4 | 0 | 0.73269 | 0.00000000 | 0.00000000 to 0.48989 |
| SecJudge | 0.0 to 0.1 | 2 | 0 | 0.02325 | 0.00000000 | 0.00000000 to 0.65762 |
| SecJudge | 0.2 to 0.3 | 47 | 1 | 0.27545 | 0.02128 | 0.00376577 to 0.11113 |
| SecJudge | 0.3 to 0.4 | 40 | 0 | 0.30917 | 0.00000000 | 0.00000000 to 0.08762 |
| SecJudge | 0.4 to 0.5 | 1 | 0 | 0.48549 | 0.00000000 | 0.00000000 to 0.79345 |
| SecJudge | 0.6 to 0.7 | 515 | 19 | 0.68252 | 0.03689 | 0.02374 to 0.05690 |
| SecJudge | 0.7 to 0.8 | 391 | 7 | 0.71197 | 0.01790 | 0.00869859 to 0.03649 |
| SecJudge | 0.8 to 0.9 | 1,192 | 164 | 0.87564 | 0.13758 | 0.11919 to 0.15831 |
| SecJudge | 0.9 to 1.0 | 1,629 | 245 | 0.93402 | 0.15040 | 0.13386 to 0.16858 |
outputs/curves/curves-s2.json :: arms.<arm>.calibration.buckets; except the log height of each count bar, which is a scale choice and is stated in the plot.Notes on individual models
- open-jev-qwen-27b ranks first at the budget but blocks only 87 unsafe cases at its shipped threshold (F1 0.33206), so it needs its own threshold before it can be deployed. Its cascade has not been run at that threshold.
- OpenJev has the lowest benign false-positive rate of the 3 models run on real coding traffic (0.01944 of benign events flagged, against 0.33117 for Jev 1.13.0) and the highest recall on proof-verified compromises of the 4 models run on them (0.5778; Intent).
- Jev 1.13.0 is a hosted API, billed at $0.042 per million input tokens; it is third at the budget.
- gemma-4-26B-A4B-it is the judge’s own weights read out as a typed decision instead of a chat answer. jevify-gemma4-26b-a4b is a LoRA fine-tune of those weights and scores below them both at the budget and as shipped.
- kev-9b answers
allowon 96.9% of cases as shipped, which gives it the lowest shipped F1 on the board; bespoke-nimble-9b answersconfirmon 82.8%, and a confirm is a miss on the block-only lens. - SecJudge blocks 98.4% of benign cases as shipped, which leaves its F1 +0.00221 above a policy that blocks everything. At the budget it makes no true block at all.
- decider-2b appears only in the re-thresholding pass. It has no board scorecard of its own; its shipped figure is the one that pass recorded.
- Von 1.0.1, a CPU classifier, was scored only on the 200-case pilot, a different corpus, so it is not in the table. There it beats a block-everything policy by +0.05241 F1 by hard-blocking 50 of 70 benign cases, at a median of 38.8s to 44.4s per case.
What the ranking rests on
| Assumption | What it rests on | What breaks it |
|---|---|---|
| The budget 0.00384502 is the right false-block price. | It is OpenJev’s own realised block false-positive rate on this corpus, so it is a shipped operating point rather than a round number. At 3,381 benign cases it allows 13 false blocks. | A deployment that tolerates more false blocks reorders the table: the ranking is a ranking at one budget, and the precision-recall curves show each model at every other budget it could be run at. |
| open-jev-qwen-27b is the best model. | F1 0.73876 at the budget, against the next model’s 0.71162, with a bootstrap 95% interval on F1 of 0.70058 to 0.77309. | The two intervals overlap, so this is a lead and not a separation. A rerun on a different case sample can reorder the top two. |
| Recall on this corpus means recall in deployment. | 382 of the 436 positives come from one source dataset, lihaonan0716/mcphunt-agent-traces, which is 87.61% of them. | A deployment whose attack mix does not look like that source is not described by these recall figures. The per-source matrix on Method prints each model against each source with the positive count behind every rate. |
| A model’s F1 at the budget is a property of the model. | Every figure at the budget is read on one variable, P(block) under aggregation definition A, at one threshold rule, so calibration is held fixed across the table. | The shipped column mixes each model with its own calibration, which is why it is a column and not the ranking. |
| Accuracy is a useful summary here. | It is not, at this prevalence: deciding allow on every case scores 0.88577. The accuracy figure is published beside that baseline and as a signed case count. | Any accuracy claim that does not carry the all-allow baseline is uninformative on a corpus that is 88.58% benign. |
| The curves and the AUC column agree. | 12 of the 12 AUCs computed for the curves were compared against the figure the published re-mining pass recorded for the same arm, variable and definition, at a tolerance of 5e-15; 0 disagree. | A rescore that moved an AUC without moving its curve would abort the pass that writes them, so the two cannot drift apart silently. |
evaluation-only never-train
Nothing derived here is approved for training, synthetic generation, teacher context, distillation or redistribution. The 13 source datasets are public and linked on reproduce; the analysis artifacts behind each number are held privately and are available on request.