Which small model to put in front of the judge

A tool-call guard here is a cascade: deterministic rules, then a small model, then the Gemma 4 judge on the cases the small model is unsure of. Each small model answers allow, confirm or block for one agent tool call. This page ranks the candidates for the middle tier on the Broad comparison corpus: 3,817 scorable cases, 436 of them unsafe and 3,381 benign. Every model is held to one shared budgetOne block false-positive rate every model is held to. Each model’s threshold is set so it makes no more false blocks than OpenJev makes on this corpus as shipped, then its recall and F1 are read at that point..

The answer

Recommended stack
rules → OpenJev → Gemma 4 judge, two-sided routingThe small model settles a case alone when it is confident either way; only the uncertain band between the allow and block thresholds goes to the judge. at an allow threshold of 0.30, with escalate-on-confirmWhen the rule engine answers confirm, that answer becomes a floor and the later tiers still run, so a case the small model or the judge would block ends as a block.: block-only F1 0.75173 at block FPR 0.00414 (14 false blocks), with the judge called on 15.95% of cases. The judge alone scores 0.71248 at block FPR 0.00680. OpenJev weights are CC BY-NC 4.0: non-commercial, attribution required, contact the authors for commercial use.
The stack as it runs today
The same stack as it runs today, with a short-circuitAny rule-engine answer other than allow ends the cascade. This is how the cascade runs today; an advisory rule confirm therefore caps a case a later tier would block. on any rule-engine answer, scores 0.73773. The two differ only on the 8 cases where an advisory rule confirm ended the cascade before a later tier could block: 263 true blocks become 271, at the same 14 false blocks and the same judge-call rate. The recommendation is the escalating version; the cost panels on Deploy are measured on the version that runs today.
Best single model at the shared budget
open-jev-qwen-27b, F1 0.73876 (interval 0.70058–0.77309), then OpenJev at 0.71162 (0.67169–0.75034). The intervals overlap, so this is a lead and not a separation. open-jev-qwen-27b is licensed apache-2.0. Its budget threshold was set on these cases; chosen out of fold it scores 0.75138 but at block FPR 0.00473, over the budget. At its shipped threshold it scores 0.33206, and in front of the judge its best measured cascade is 0.59351, below the judge alone. No cascade has been run with it at the budget threshold, which is why the recommended stack uses OpenJev.
At production traffic
On the Production-weighted corpus (0.90% unsafe) the judge costs more than it adds: rules → OpenJev with no judge scores 0.21285 against 0.08215 for the full cascade and 0.10980 for the judge alone. See Deploy.

Every model at the shared budget

One ranking. Each model’s threshold is set so it makes no more than 13 false blocks, and its block-only F1F1 counting only a hard block as a catch. A confirm counts as a miss, because a confirm does not stop the tool call., recall and precision are read there. The shipped column is the same model at the threshold it ships with, which mixes the model with its calibration; it is printed for reference and does not rank anything.

Every model at block FPR ≤ 0.00384502 (at most 13 false blocks in 3,381 benign cases) on the Broad comparison, 436 unsafe cases. Ranked by F1 at that budget; the interval is a family bootstrap with the threshold held fixed. Each threshold was set on these same cases, so the budget figures are an upper bound on what the model does at a threshold fixed in advance. Source: outputs/curves/curves-s2.json.
RankModelLicenseF1 at the budget
95% interval
recallprecisiontrue / false blocksAUCblock-only F1 as shippedparameters
1open-jev-qwen-27bapache-2.00.73876
0.70058–0.77309
0.603210.95290263 / 130.948030.3320627B (nominal)
2OpenJevCC BY-NC 4.0 (non-commercial)0.71162
0.67169–0.75034
0.568810.95019248 / 130.974540.70231not recorded
3Jev 1.13.0commercial API, $0.042/M input tokens0.58360
0.53443–0.62829
0.424310.93434185 / 130.939420.54153not recorded
4gemma-4-26B-A4B-itapache-2.00.32463
0.27206–0.37113
0.199540.8700087 / 130.885500.4749625.8B (counted)
5jevify-gemma4-26b-a4bgemma0.29278
0.24231–0.34019
0.176610.8555677 / 130.879890.0492225.8B (counted)
6open-jev-qwen-9bapache-2.00.24609
0.19706–0.29769
0.144500.8289563 / 130.782690.175519B (nominal)
7DiffusionGemmaapache-2.00.22530
0.17923–0.27495
0.130730.8142957 / 130.777290.2679226B (nominal)
8kev-9bapache-2.00.19316
0.14729–0.23886
0.110090.7868948 / 130.826220.018149B (nominal)
9open-jev-qwen-2bapache-2.00.08120
0.04785–0.11861
0.043580.5937519 / 130.553520.171952B (nominal)
10decider-2b re-mining onlynot recorded0.07296
0.04184–0.10856
0.038990.5666717 / 130.770400.108432B (nominal)
11bespoke-nimble-9bapache-2.00.00887
0.00000–0.02278
0.004590.133332 / 130.896300.215699B (nominal)
12SecJudgeapache-2.00.00000
0.00000–0.00000
0.000000.000000 / 10.653240.20724396M (counted)
refGemma 4 judge referenceapache-2.0 weights, served via Bedrock (paid service)—————0.71248not recorded

Every ranked model answered the same question format, Q2The question format every ranked model answered: one disposition choice (allow / confirm / block) plus two scores., except SecJudge, a classifier that takes one serialised string and answers no question at all. Question format moves F1 about as much as the choice of model does, which is why the ranking holds it fixed; the measurement is on Method.

Precision, recall and F1 per model at the shared budget

precisionrecallF1
Precision, recall and F1 per model at the shared budgetOne row per model, ordered by F1. Three dots on a shared 0 to 1 axis: precision, recall and F1, each at the shared block-FPR budget 0.00384502. Precision and recall carry a Wilson 95% interval and F1 a bootstrap 95% interval. Every model at one block false-positive budget, 0.00384502, on one ranking variable. Whiskers: Wilson 95% on precision and on recall, percentile bootstrap 95% on F1. Rows are sorted by F1, highest first. 0 0.25 0.5 0.75 1 open-jev-qwen-27b open-jev-qwen-27b · precision 0.95290, Wilson 95% 0.92109 to 0.97227 open-jev-qwen-27b · recall 0.60321, Wilson 95% 0.55658 to 0.64804 open-jev-qwen-27b · F1 0.73876, bootstrap 95% 0.70058 to 0.77309 0.73876 OpenJev OpenJev · precision 0.95019, Wilson 95% 0.91666 to 0.97066 OpenJev · recall 0.56881, Wilson 95% 0.52192 to 0.61449 OpenJev · F1 0.71162, bootstrap 95% 0.67169 to 0.75034 0.71162 Jev 1.13.0 Jev 1.13.0 · precision 0.93434, Wilson 95% 0.89092 to 0.96123 Jev 1.13.0 · recall 0.42431, Wilson 95% 0.37878 to 0.47117 Jev 1.13.0 · F1 0.58360, bootstrap 95% 0.53443 to 0.62829 0.58360 gemma-4-26B-A4B-it gemma-4-26B-A4B-it · precision 0.87000, Wilson 95% 0.79020 to 0.92243 gemma-4-26B-A4B-it · recall 0.19954, Wilson 95% 0.16472 to 0.23961 gemma-4-26B-A4B-it · F1 0.32463, bootstrap 95% 0.27206 to 0.37113 0.32463 jevify-gemma4-26b-a4b jevify-gemma4-26b-a4b · precision 0.85556, Wilson 95% 0.76840 to 0.91360 jevify-gemma4-26b-a4b · recall 0.17661, Wilson 95% 0.14368 to 0.21518 jevify-gemma4-26b-a4b · F1 0.29278, bootstrap 95% 0.24231 to 0.34019 0.29278 open-jev-qwen-9b open-jev-qwen-9b · precision 0.82895, Wilson 95% 0.72902 to 0.89722 open-jev-qwen-9b · recall 0.14450, Wilson 95% 0.11460 to 0.18060 open-jev-qwen-9b · F1 0.24609, bootstrap 95% 0.19706 to 0.29769 0.24609 DiffusionGemma DiffusionGemma · precision 0.81429, Wilson 95% 0.70774 to 0.88813 DiffusionGemma · recall 0.13073, Wilson 95% 0.10229 to 0.16563 DiffusionGemma · F1 0.22530, bootstrap 95% 0.17923 to 0.27495 0.22530 kev-9b kev-9b · precision 0.78689, Wilson 95% 0.66878 to 0.87100 kev-9b · recall 0.11009, Wilson 95% 0.08405 to 0.14295 kev-9b · F1 0.19316, bootstrap 95% 0.14729 to 0.23886 0.19316 open-jev-qwen-2b open-jev-qwen-2b · precision 0.59375, Wilson 95% 0.42260 to 0.74480 open-jev-qwen-2b · recall 0.04358, Wilson 95% 0.02807 to 0.06706 open-jev-qwen-2b · F1 0.08120, bootstrap 95% 0.04785 to 0.11861 0.08120 decider-2b decider-2b · precision 0.56667, Wilson 95% 0.39197 to 0.72623 decider-2b · recall 0.03899, Wilson 95% 0.02448 to 0.06155 decider-2b · F1 0.07296, bootstrap 95% 0.04184 to 0.10856 0.07296 bespoke-nimble-9b bespoke-nimble-9b · precision 0.13333, Wilson 95% 0.03736 to 0.37882 bespoke-nimble-9b · recall 0.00458716, Wilson 95% 0.00125887 to 0.01657 bespoke-nimble-9b · F1 0.00886918, bootstrap 95% 0.00000000 to 0.02278 0.00887 SecJudge SecJudge · precision 0.00000000, Wilson 95% 0.00000000 to 0.79345 SecJudge · recall 0.00000000, Wilson 95% 0.00000000 to 0.00873374 SecJudge · F1 0.00000000, bootstrap 95% 0.00000000 to 0.00000000 0.00000 precision, recall and F1 over 436 positives and 3,381 benign cases
Data table
Sorted by F1 at the shared budget, highest first.
Modelprecisionprecision, Wilson 95%recallrecall, Wilson 95%F1F1, bootstrap 95%
open-jev-qwen-27b0.952900.92109 to 0.972270.603210.55658 to 0.648040.738760.70058 to 0.77309
OpenJev0.950190.91666 to 0.970660.568810.52192 to 0.614490.711620.67169 to 0.75034
Jev 1.13.00.934340.89092 to 0.961230.424310.37878 to 0.471170.583600.53443 to 0.62829
gemma-4-26B-A4B-it0.870000.79020 to 0.922430.199540.16472 to 0.239610.324630.27206 to 0.37113
jevify-gemma4-26b-a4b0.855560.76840 to 0.913600.176610.14368 to 0.215180.292780.24231 to 0.34019
open-jev-qwen-9b0.828950.72902 to 0.897220.144500.11460 to 0.180600.246090.19706 to 0.29769
DiffusionGemma0.814290.70774 to 0.888130.130730.10229 to 0.165630.225300.17923 to 0.27495
kev-9b0.786890.66878 to 0.871000.110090.08405 to 0.142950.193160.14729 to 0.23886
open-jev-qwen-2b0.593750.42260 to 0.744800.043580.02807 to 0.067060.081200.04785 to 0.11861
decider-2b0.566670.39197 to 0.726230.038990.02448 to 0.061550.072960.04184 to 0.10856
bespoke-nimble-9b0.133330.03736 to 0.378820.004587160.00125887 to 0.016570.008869180.00000000 to 0.02278
SecJudge0.000000000.00000000 to 0.793450.000000000.00000000 to 0.008733740.000000000.00000000 to 0.00000000
Each model spends the same false-block allowance, so the three dots on a row are the same decision read three ways. The interval on F1 is a bootstrap because F1 is not a single proportion; the intervals on precision and recall are Wilson because each is. Source: outputs/curves/curves-s2.json :: arms.<arm>.at_budget.

Every model at every other budget

The table is one point on each model’s precision-recall curve. A deployment that tolerates more false blocks reads a different point, and the order can change.

Precision against recall per model, with prevalence drawn

Prevalence 0.11423 is the horizontal baseline on every panel.

the model's precision-recall curveits point at the shared budget
Precision against recall per model, with prevalence and the shared budget point12 panels. Each plots precision against recall over the 3,817 scorable cases, draws the corpus prevalence 0.11423 as a horizontal baseline, and marks the model’s point at the shared block-FPR budget 0.00384502 as a dot. Precision on the vertical axis, recall on the horizontal, both linear from 0 to 1, per case, on P(block) under definition A. Dashed horizontal rule: the corpus prevalence 0.11423. That is the precision a policy that blocks at random reaches, so a curve is only above it where it is doing work. Dot: the model’s point at the shared block-FPR budget 0.00384502. Panels are ordered by F1 at the budget, highest first. open-jev-qwen-27b open-jev-qwen-27b at the budget: precision 0.95290, recall 0.60321, F1 0.73876 F1 0.73876 0 0.5 1 OpenJev OpenJev at the budget: precision 0.95019, recall 0.56881, F1 0.71162 F1 0.71162 0 0.5 1 Jev 1.13.0 Jev 1.13.0 at the budget: precision 0.93434, recall 0.42431, F1 0.58360 F1 0.58360 0 0.5 1 gemma-4-26B-A4B-it gemma-4-26B-A4B-it at the budget: precision 0.87000, recall 0.19954, F1 0.32463 F1 0.32463 0 0.5 1 jevify-gemma4-26b-a4b jevify-gemma4-26b-a4b at the budget: precision 0.85556, recall 0.17661, F1 0.29278 F1 0.29278 0 0.5 1 open-jev-qwen-9b open-jev-qwen-9b at the budget: precision 0.82895, recall 0.14450, F1 0.24609 F1 0.24609 0 0.5 1 DiffusionGemma DiffusionGemma at the budget: precision 0.81429, recall 0.13073, F1 0.22530 F1 0.22530 0 0.5 1 kev-9b kev-9b at the budget: precision 0.78689, recall 0.11009, F1 0.19316 F1 0.19316 0 0.5 1 open-jev-qwen-2b open-jev-qwen-2b at the budget: precision 0.59375, recall 0.04358, F1 0.08120 F1 0.08120 0 0.5 1 decider-2b decider-2b at the budget: precision 0.56667, recall 0.03899, F1 0.07296 F1 0.07296 0 0.5 1 bespoke-nimble-9b bespoke-nimble-9b at the budget: precision 0.13333, recall 0.00458716, F1 0.00886918 F1 0.00887 0 0.5 1 SecJudge SecJudge at the budget: precision 0.00000000, recall 0.00000000, F1 0.00000000 F1 0.00000 0 0.5 1
Data table
Every plotted model at the shared budget, ordered by F1, highest first. The interval method is named in each column header.
Modelpoints plottedprecisionrecallF1precision, Wilson 95%recall, Wilson 95%F1, bootstrap 95%
open-jev-qwen-27b4110.952900.603210.738760.92109 to 0.972270.55658 to 0.648040.70058 to 0.77309
OpenJev3030.950190.568810.711620.91666 to 0.970660.52192 to 0.614490.67169 to 0.75034
Jev 1.13.0930.934340.424310.583600.89092 to 0.961230.37878 to 0.471170.53443 to 0.62829
gemma-4-26B-A4B-it2650.870000.199540.324630.79020 to 0.922430.16472 to 0.239610.27206 to 0.37113
jevify-gemma4-26b-a4b2790.855560.176610.292780.76840 to 0.913600.14368 to 0.215180.24231 to 0.34019
open-jev-qwen-9b2050.828950.144500.246090.72902 to 0.897220.11460 to 0.180600.19706 to 0.29769
DiffusionGemma2380.814290.130730.225300.70774 to 0.888130.10229 to 0.165630.17923 to 0.27495
kev-9b2050.786890.110090.193160.66878 to 0.871000.08405 to 0.142950.14729 to 0.23886
open-jev-qwen-2b1520.593750.043580.081200.42260 to 0.744800.02807 to 0.067060.04785 to 0.11861
decider-2b1630.566670.038990.072960.39197 to 0.726230.02448 to 0.061550.04184 to 0.10856
bespoke-nimble-9b1450.133330.004587160.008869180.03736 to 0.378820.00125887 to 0.016570.00000000 to 0.02278
SecJudge980.000000000.000000000.000000000.00000000 to 0.793450.00000000 to 0.008733740.00000000 to 0.00000000
Precision and recall carry a Wilson 95% interval on their own count; F1 carries a percentile bootstrap 95% over 2,000 resamples of families with the threshold held fixed. On this corpus strata.split_group is unique per case, so its 3,817 families are its 3,817 scorable cases and the family bootstrap is a case bootstrap. Source: outputs/curves/curves-s2.json :: arms.<arm>.pr, arms.<arm>.at_budget and corpus.prevalence.

Size against score

Parameter counts are counted from the served weights where the artifacts record them and nominal (read from the base model’s name) otherwise; 2 (OpenJev and Jev 1.13.0) carry neither and are not plotted.

Parameters against F1 at the shared budget

3 counted figures and 7 nominal ones.

counted parameter figure
Parameters against F1 at the shared budget10 banded models. Horizontal axis: parameter count on a log scale, with the 3B and 6B band boundaries drawn. Vertical axis: F1 at the shared block-FPR budget. A filled point is a counted parameter figure and a hollow point is the nominal size of the base model the serving record names. Parameter count on a log scale against F1 at the shared block-FPR budget 0.00384502. 3 of these 10 models have a counted parameter figure. The other 7 are banded on the nominal size of the base model their serving record names. A hollow point is a nominal figure, so its horizontal position is a model name rather than a measurement. Dashed vertical rules: the 3B and 6B band boundaries. 2 further models carry no parameter figure of either kind and are not on this plot: OpenJev, Jev 1.13.0. 0 0.25 0.5 0.75 1 3B 6B 250M 1B 10B 30B SecJudge: 395,836,421 parameters (counted), F1 0.00000000 at the shared budget SecJudge open-jev-qwen-2b: 2,000,000,000 parameters (nominal), F1 0.08120 at the shared budget open-jev-qwen-2b decider-2b: 2,000,000,000 parameters (nominal), F1 0.07296 at the shared budget decider-2b open-jev-qwen-9b: 9,000,000,000 parameters (nominal), F1 0.24609 at the shared budget open-jev-qwen-9b kev-9b: 9,000,000,000 parameters (nominal), F1 0.19316 at the shared budget kev-9b bespoke-nimble-9b: 9,000,000,000 parameters (nominal), F1 0.00886918 at the shared budget bespoke-nimble-9b gemma-4-26B-A4B-it: 25,805,936,206 parameters (counted), F1 0.32463 at the shared budget gemma-4-26B-A4B-it jevify-gemma4-26b-a4b: 25,805,936,206 parameters (counted), F1 0.29278 at the shared budget jevify-gemma4-26b-a4b DiffusionGemma: 26,000,000,000 parameters (nominal), F1 0.22530 at the shared budget DiffusionGemma open-jev-qwen-27b: 27,000,000,000 parameters (nominal), F1 0.73876 at the shared budget open-jev-qwen-27b counted or nominal parameters, log scale
Data table
Sorted by parameter count, smallest first.
ModelparametersbasisbandF1 at the shared budget
SecJudge395,836,421countedunder 3B0.00000000
open-jev-qwen-2b2,000,000,000nominalunder 3B0.08120
decider-2b2,000,000,000nominalunder 3B0.07296
open-jev-qwen-9b9,000,000,000nominal6B and up0.24609
kev-9b9,000,000,000nominal6B and up0.19316
bespoke-nimble-9b9,000,000,000nominal6B and up0.00886918
gemma-4-26B-A4B-it25,805,936,206counted6B and up0.32463
jevify-gemma4-26b-a4b25,805,936,206counted6B and up0.29278
DiffusionGemma26,000,000,000nominal6B and up0.22530
open-jev-qwen-27b27,000,000,000nominal6B and up0.73876
A band boundary crossed by a nominal figure is a boundary crossed by a model name. 7 of the 10 points are nominal, so the shape of this plot is partly a naming convention. Source: each arm's serving record, outputs/gemma4jev/artifacts/out/jevify-weight-diff.json, and outputs/curves/curves-s2.json :: arms.<arm>.at_budget.f1; except the log position of each point, and the hollow fill that marks a nominal figure.

Does the score mean what it says

A threshold is only portable if the probability behind it is calibrated. Each panel buckets the per-case score and plots the share of the bucket that is unsafe, with the case count beside it.

Predicted probability against observed positive rate, per model

10 equal-width buckets of the per-case maximum P(block), with each bucket’s case count printed.

observed positive ratecases in the bucket, log scale
Predicted probability against observed positive rate, per model12 panels, 10 equal-width buckets of the per-case maximum P(block) each. The dot is the share of that bucket’s cases that are grade A or B positives, with a Wilson 95% interval. The grey bar behind each bucket is its case count on a log scale, and the count is printed under the axis, so a bucket holding a handful of cases is visible as one. Horizontal axis: the per-case maximum P(block), in 10 equal-width buckets. Vertical axis: the share of that bucket that is a grade A or B positive. A calibrated model follows the dashed diagonal. Whiskers are Wilson 95% on the bucket’s own count. Grey bar: the bucket’s case count on a log scale, full height at 3,817 cases. The count is printed under each bucket. A bucket with no cases carries no dot and no interval. Panels are ordered by F1 at the shared budget, highest first. open-jev-qwen-27b 0 0.5 1 open-jev-qwen-27b · bucket [0.0, 0.1]: 3,390 cases, 84 positive, observed rate 0.02478, mean predicted 0.01977 3,390 open-jev-qwen-27b · bucket [0.1, 0.2]: 166 cases, 102 positive, observed rate 0.61446, mean predicted 0.14237 166 open-jev-qwen-27b · bucket [0.2, 0.3]: 109 cases, 100 positive, observed rate 0.91743, mean predicted 0.24438 109 open-jev-qwen-27b · bucket [0.3, 0.4]: 85 cases, 84 positive, observed rate 0.98824, mean predicted 0.35248 85 open-jev-qwen-27b · bucket [0.4, 0.5]: 30 cases, 30 positive, observed rate 1, mean predicted 0.44028 30 open-jev-qwen-27b · bucket [0.5, 0.6]: 22 cases, 21 positive, observed rate 0.95455, mean predicted 0.54063 22 open-jev-qwen-27b · bucket [0.6, 0.7]: 4 cases, 4 positive, observed rate 1, mean predicted 0.64929 4 open-jev-qwen-27b · bucket [0.7, 0.8]: 3 cases, 3 positive, observed rate 1, mean predicted 0.77815 3 open-jev-qwen-27b · bucket [0.8, 0.9]: 6 cases, 6 positive, observed rate 1, mean predicted 0.86237 6 open-jev-qwen-27b · bucket [0.9, 1.0]: 2 cases, 2 positive, observed rate 1, mean predicted 0.94717 2 cases per bucket OpenJev 0 0.5 1 OpenJev · bucket [0.0, 0.1]: 3,372 cases, 81 positive, observed rate 0.02402, mean predicted 0.01313 3,372 OpenJev · bucket [0.1, 0.2]: 88 cases, 48 positive, observed rate 0.54545, mean predicted 0.14455 88 OpenJev · bucket [0.2, 0.3]: 64 cases, 32 positive, observed rate 0.50000, mean predicted 0.25181 64 OpenJev · bucket [0.3, 0.4]: 39 cases, 33 positive, observed rate 0.84615, mean predicted 0.34620 39 OpenJev · bucket [0.4, 0.5]: 49 cases, 43 positive, observed rate 0.87755, mean predicted 0.45022 49 OpenJev · bucket [0.5, 0.6]: 42 cases, 38 positive, observed rate 0.90476, mean predicted 0.54498 42 OpenJev · bucket [0.6, 0.7]: 47 cases, 46 positive, observed rate 0.97872, mean predicted 0.65070 47 OpenJev · bucket [0.7, 0.8]: 34 cases, 34 positive, observed rate 1, mean predicted 0.75730 34 OpenJev · bucket [0.8, 0.9]: 58 cases, 57 positive, observed rate 0.98276, mean predicted 0.84549 58 OpenJev · bucket [0.9, 1.0]: 24 cases, 24 positive, observed rate 1, mean predicted 0.94472 24 cases per bucket Jev 1.13.0 0 0.5 1 Jev 1.13.0 · bucket [0.0, 0.1]: 3,426 cases, 102 positive, observed rate 0.02977, mean predicted 0.00514302 3,426 Jev 1.13.0 · bucket [0.1, 0.2]: 87 cases, 61 positive, observed rate 0.70115, mean predicted 0.14517 87 Jev 1.13.0 · bucket [0.2, 0.3]: 73 cases, 59 positive, observed rate 0.80822, mean predicted 0.24096 73 Jev 1.13.0 · bucket [0.3, 0.4]: 68 cases, 59 positive, observed rate 0.86765, mean predicted 0.34588 68 Jev 1.13.0 · bucket [0.4, 0.5]: 58 cases, 52 positive, observed rate 0.89655, mean predicted 0.44121 58 Jev 1.13.0 · bucket [0.5, 0.6]: 41 cases, 41 positive, observed rate 1, mean predicted 0.54415 41 Jev 1.13.0 · bucket [0.6, 0.7]: 30 cases, 29 positive, observed rate 0.96667, mean predicted 0.64133 30 Jev 1.13.0 · bucket [0.7, 0.8]: 16 cases, 15 positive, observed rate 0.93750, mean predicted 0.73188 16 Jev 1.13.0 · bucket [0.8, 0.9]: 8 cases, 8 positive, observed rate 1, mean predicted 0.84125 8 Jev 1.13.0 · bucket [0.9, 1.0]: 10 cases, 10 positive, observed rate 1, mean predicted 0.95000 10 cases per bucket gemma-4-26B-A4B-it 0 0.5 1 gemma-4-26B-A4B-it · bucket [0.0, 0.1]: 3,003 cases, 112 positive, observed rate 0.03730, mean predicted 0.05352 3,003 gemma-4-26B-A4B-it · bucket [0.1, 0.2]: 529 cases, 125 positive, observed rate 0.23629, mean predicted 0.12817 529 gemma-4-26B-A4B-it · bucket [0.2, 0.3]: 65 cases, 34 positive, observed rate 0.52308, mean predicted 0.24006 65 gemma-4-26B-A4B-it · bucket [0.3, 0.4]: 31 cases, 16 positive, observed rate 0.51613, mean predicted 0.34656 31 gemma-4-26B-A4B-it · bucket [0.4, 0.5]: 38 cases, 20 positive, observed rate 0.52632, mean predicted 0.44194 38 gemma-4-26B-A4B-it · bucket [0.5, 0.6]: 31 cases, 26 positive, observed rate 0.83871, mean predicted 0.54915 31 gemma-4-26B-A4B-it · bucket [0.6, 0.7]: 42 cases, 35 positive, observed rate 0.83333, mean predicted 0.65084 42 gemma-4-26B-A4B-it · bucket [0.7, 0.8]: 35 cases, 29 positive, observed rate 0.82857, mean predicted 0.74933 35 gemma-4-26B-A4B-it · bucket [0.8, 0.9]: 35 cases, 31 positive, observed rate 0.88571, mean predicted 0.83584 35 gemma-4-26B-A4B-it · bucket [0.9, 1.0]: 8 cases, 8 positive, observed rate 1, mean predicted 0.92490 8 cases per bucket jevify-gemma4-26b-a4b 0 0.5 1 jevify-gemma4-26b-a4b · bucket [0.0, 0.1]: 3,773 cases, 396 positive, observed rate 0.10496, mean predicted 0.00292870 3,773 jevify-gemma4-26b-a4b · bucket [0.1, 0.2]: 20 cases, 17 positive, observed rate 0.85000, mean predicted 0.12455 20 jevify-gemma4-26b-a4b · bucket [0.2, 0.3]: 10 cases, 10 positive, observed rate 1, mean predicted 0.24279 10 jevify-gemma4-26b-a4b · bucket [0.3, 0.4]: 3 cases, 2 positive, observed rate 0.66667, mean predicted 0.32102 3 jevify-gemma4-26b-a4b · bucket [0.4, 0.5]: 5 cases, 5 positive, observed rate 1, mean predicted 0.43032 5 jevify-gemma4-26b-a4b · bucket [0.5, 0.6]: 1 cases, 1 positive, observed rate 1, mean predicted 0.56031 1 jevify-gemma4-26b-a4b · bucket [0.6, 0.7]: 1 cases, 1 positive, observed rate 1, mean predicted 0.67752 1 jevify-gemma4-26b-a4b · bucket [0.7, 0.8]: 1 cases, 1 positive, observed rate 1, mean predicted 0.79245 1 0 jevify-gemma4-26b-a4b · bucket [0.9, 1.0]: 3 cases, 3 positive, observed rate 1, mean predicted 0.95399 3 cases per bucket open-jev-qwen-9b 0 0.5 1 open-jev-qwen-9b · bucket [0.0, 0.1]: 3,679 cases, 341 positive, observed rate 0.09269, mean predicted 0.02260 3,679 open-jev-qwen-9b · bucket [0.1, 0.2]: 59 cases, 32 positive, observed rate 0.54237, mean predicted 0.13512 59 open-jev-qwen-9b · bucket [0.2, 0.3]: 17 cases, 13 positive, observed rate 0.76471, mean predicted 0.24563 17 open-jev-qwen-9b · bucket [0.3, 0.4]: 10 cases, 8 positive, observed rate 0.80000, mean predicted 0.33785 10 open-jev-qwen-9b · bucket [0.4, 0.5]: 8 cases, 4 positive, observed rate 0.50000, mean predicted 0.44196 8 open-jev-qwen-9b · bucket [0.5, 0.6]: 8 cases, 7 positive, observed rate 0.87500, mean predicted 0.54878 8 open-jev-qwen-9b · bucket [0.6, 0.7]: 3 cases, 2 positive, observed rate 0.66667, mean predicted 0.65367 3 open-jev-qwen-9b · bucket [0.7, 0.8]: 11 cases, 10 positive, observed rate 0.90909, mean predicted 0.77306 11 open-jev-qwen-9b · bucket [0.8, 0.9]: 5 cases, 5 positive, observed rate 1, mean predicted 0.84864 5 open-jev-qwen-9b · bucket [0.9, 1.0]: 17 cases, 14 positive, observed rate 0.82353, mean predicted 0.95487 17 cases per bucket DiffusionGemma 0 0.5 1 DiffusionGemma · bucket [0.0, 0.1]: 3,547 cases, 263 positive, observed rate 0.07415, mean predicted 0.00826698 3,547 DiffusionGemma · bucket [0.1, 0.2]: 81 cases, 42 positive, observed rate 0.51852, mean predicted 0.14096 81 DiffusionGemma · bucket [0.2, 0.3]: 52 cases, 29 positive, observed rate 0.55769, mean predicted 0.25021 52 DiffusionGemma · bucket [0.3, 0.4]: 40 cases, 28 positive, observed rate 0.70000, mean predicted 0.34126 40 DiffusionGemma · bucket [0.4, 0.5]: 21 cases, 15 positive, observed rate 0.71429, mean predicted 0.45379 21 DiffusionGemma · bucket [0.5, 0.6]: 18 cases, 13 positive, observed rate 0.72222, mean predicted 0.53993 18 DiffusionGemma · bucket [0.6, 0.7]: 13 cases, 8 positive, observed rate 0.61538, mean predicted 0.66025 13 DiffusionGemma · bucket [0.7, 0.8]: 24 cases, 18 positive, observed rate 0.75000, mean predicted 0.75377 24 DiffusionGemma · bucket [0.8, 0.9]: 8 cases, 7 positive, observed rate 0.87500, mean predicted 0.83024 8 DiffusionGemma · bucket [0.9, 1.0]: 13 cases, 13 positive, observed rate 1, mean predicted 0.96392 13 cases per bucket kev-9b 0 0.5 1 kev-9b · bucket [0.0, 0.1]: 3,392 cases, 264 positive, observed rate 0.07783, mean predicted 0.04014 3,392 kev-9b · bucket [0.1, 0.2]: 371 cases, 130 positive, observed rate 0.35040, mean predicted 0.12893 371 kev-9b · bucket [0.2, 0.3]: 44 cases, 36 positive, observed rate 0.81818, mean predicted 0.22838 44 kev-9b · bucket [0.3, 0.4]: 4 cases, 1 positive, observed rate 0.25000, mean predicted 0.33643 4 kev-9b · bucket [0.4, 0.5]: 1 cases, 1 positive, observed rate 1, mean predicted 0.40080 1 kev-9b · bucket [0.5, 0.6]: 2 cases, 1 positive, observed rate 0.50000, mean predicted 0.57687 2 kev-9b · bucket [0.6, 0.7]: 2 cases, 2 positive, observed rate 1, mean predicted 0.66204 2 0 0 kev-9b · bucket [0.9, 1.0]: 1 cases, 1 positive, observed rate 1, mean predicted 0.95163 1 cases per bucket open-jev-qwen-2b 0 0.5 1 open-jev-qwen-2b · bucket [0.0, 0.1]: 3,273 cases, 337 positive, observed rate 0.10296, mean predicted 0.03709 3,273 open-jev-qwen-2b · bucket [0.1, 0.2]: 208 cases, 31 positive, observed rate 0.14904, mean predicted 0.13822 208 open-jev-qwen-2b · bucket [0.2, 0.3]: 59 cases, 7 positive, observed rate 0.11864, mean predicted 0.24581 59 open-jev-qwen-2b · bucket [0.3, 0.4]: 40 cases, 3 positive, observed rate 0.07500, mean predicted 0.35053 40 open-jev-qwen-2b · bucket [0.4, 0.5]: 40 cases, 3 positive, observed rate 0.07500, mean predicted 0.44373 40 open-jev-qwen-2b · bucket [0.5, 0.6]: 47 cases, 7 positive, observed rate 0.14894, mean predicted 0.54262 47 open-jev-qwen-2b · bucket [0.6, 0.7]: 45 cases, 8 positive, observed rate 0.17778, mean predicted 0.65060 45 open-jev-qwen-2b · bucket [0.7, 0.8]: 35 cases, 8 positive, observed rate 0.22857, mean predicted 0.75057 35 open-jev-qwen-2b · bucket [0.8, 0.9]: 28 cases, 10 positive, observed rate 0.35714, mean predicted 0.85511 28 open-jev-qwen-2b · bucket [0.9, 1.0]: 42 cases, 22 positive, observed rate 0.52381, mean predicted 0.94761 42 cases per bucket decider-2b 0 0.5 1 decider-2b · bucket [0.0, 0.1]: 3,267 cases, 251 positive, observed rate 0.07683, mean predicted 0.04746 3,267 decider-2b · bucket [0.1, 0.2]: 379 cases, 111 positive, observed rate 0.29288, mean predicted 0.13346 379 decider-2b · bucket [0.2, 0.3]: 72 cases, 27 positive, observed rate 0.37500, mean predicted 0.24231 72 decider-2b · bucket [0.3, 0.4]: 30 cases, 16 positive, observed rate 0.53333, mean predicted 0.34378 30 decider-2b · bucket [0.4, 0.5]: 13 cases, 6 positive, observed rate 0.46154, mean predicted 0.44099 13 decider-2b · bucket [0.5, 0.6]: 18 cases, 6 positive, observed rate 0.33333, mean predicted 0.55396 18 decider-2b · bucket [0.6, 0.7]: 16 cases, 8 positive, observed rate 0.50000, mean predicted 0.65500 16 decider-2b · bucket [0.7, 0.8]: 11 cases, 5 positive, observed rate 0.45455, mean predicted 0.73179 11 decider-2b · bucket [0.8, 0.9]: 11 cases, 6 positive, observed rate 0.54545, mean predicted 0.85790 11 0 cases per bucket bespoke-nimble-9b 0 0.5 1 bespoke-nimble-9b · bucket [0.0, 0.1]: 9 cases, 0 positive, observed rate 0.00000000, mean predicted 0.08758 9 bespoke-nimble-9b · bucket [0.1, 0.2]: 2,237 cases, 21 positive, observed rate 0.00938757, mean predicted 0.15781 2,237 bespoke-nimble-9b · bucket [0.2, 0.3]: 1,074 cases, 150 positive, observed rate 0.13966, mean predicted 0.23817 1,074 bespoke-nimble-9b · bucket [0.3, 0.4]: 325 cases, 183 positive, observed rate 0.56308, mean predicted 0.34640 325 bespoke-nimble-9b · bucket [0.4, 0.5]: 118 cases, 68 positive, observed rate 0.57627, mean predicted 0.44106 118 bespoke-nimble-9b · bucket [0.5, 0.6]: 38 cases, 12 positive, observed rate 0.31579, mean predicted 0.53819 38 bespoke-nimble-9b · bucket [0.6, 0.7]: 12 cases, 2 positive, observed rate 0.16667, mean predicted 0.63814 12 bespoke-nimble-9b · bucket [0.7, 0.8]: 4 cases, 0 positive, observed rate 0.00000000, mean predicted 0.73269 4 0 0 cases per bucket SecJudge 0 0.5 1 SecJudge · bucket [0.0, 0.1]: 2 cases, 0 positive, observed rate 0.00000000, mean predicted 0.02325 2 0 SecJudge · bucket [0.2, 0.3]: 47 cases, 1 positive, observed rate 0.02128, mean predicted 0.27545 47 SecJudge · bucket [0.3, 0.4]: 40 cases, 0 positive, observed rate 0.00000000, mean predicted 0.30917 40 SecJudge · bucket [0.4, 0.5]: 1 cases, 0 positive, observed rate 0.00000000, mean predicted 0.48549 1 0 SecJudge · bucket [0.6, 0.7]: 515 cases, 19 positive, observed rate 0.03689, mean predicted 0.68252 515 SecJudge · bucket [0.7, 0.8]: 391 cases, 7 positive, observed rate 0.01790, mean predicted 0.71197 391 SecJudge · bucket [0.8, 0.9]: 1,192 cases, 164 positive, observed rate 0.13758, mean predicted 0.87564 1,192 SecJudge · bucket [0.9, 1.0]: 1,629 cases, 245 positive, observed rate 0.15040, mean predicted 0.93402 1,629 cases per bucket
Data table
Every bucket that holds at least one case, in panel order.
Modelbucketcasespositivesmean predictedobserved rateobserved rate, Wilson 95%
open-jev-qwen-27b0.0 to 0.13,390840.019770.024780.02006 to 0.03057
open-jev-qwen-27b0.1 to 0.21661020.142370.614460.53862 to 0.68511
open-jev-qwen-27b0.2 to 0.31091000.244380.917430.85049 to 0.95595
open-jev-qwen-27b0.3 to 0.485840.352480.988240.93633 to 0.99792
open-jev-qwen-27b0.4 to 0.530300.4402810.88649 to 1
open-jev-qwen-27b0.5 to 0.622210.540630.954550.78202 to 0.99193
open-jev-qwen-27b0.6 to 0.7440.6492910.51011 to 1
open-jev-qwen-27b0.7 to 0.8330.7781510.43850 to 1
open-jev-qwen-27b0.8 to 0.9660.8623710.60967 to 1
open-jev-qwen-27b0.9 to 1.0220.9471710.34238 to 1
OpenJev0.0 to 0.13,372810.013130.024020.01937 to 0.02976
OpenJev0.1 to 0.288480.144550.545450.44170 to 0.64541
OpenJev0.2 to 0.364320.251810.500000.38102 to 0.61898
OpenJev0.3 to 0.439330.346200.846150.70271 to 0.92753
OpenJev0.4 to 0.549430.450220.877550.75756 to 0.94265
OpenJev0.5 to 0.642380.544980.904760.77935 to 0.96234
OpenJev0.6 to 0.747460.650700.978720.88887 to 0.99623
OpenJev0.7 to 0.834340.7573010.89849 to 1
OpenJev0.8 to 0.958570.845490.982760.90859 to 0.99695
OpenJev0.9 to 1.024240.9447210.86202 to 1
Jev 1.13.00.0 to 0.13,4261020.005143020.029770.02459 to 0.03601
Jev 1.13.00.1 to 0.287610.145170.701150.59813 to 0.78716
Jev 1.13.00.2 to 0.373590.240960.808220.70344 to 0.88218
Jev 1.13.00.3 to 0.468590.345880.867650.76720 to 0.92878
Jev 1.13.00.4 to 0.558520.441210.896550.79212 to 0.95172
Jev 1.13.00.5 to 0.641410.5441510.91433 to 1
Jev 1.13.00.6 to 0.730290.641330.966670.83330 to 0.99409
Jev 1.13.00.7 to 0.816150.731880.937500.71671 to 0.98888
Jev 1.13.00.8 to 0.9880.8412510.67559 to 1
Jev 1.13.00.9 to 1.010100.9500010.72247 to 1
gemma-4-26B-A4B-it0.0 to 0.13,0031120.053520.037300.03109 to 0.04469
gemma-4-26B-A4B-it0.1 to 0.25291250.128170.236290.20208 to 0.27432
gemma-4-26B-A4B-it0.2 to 0.365340.240060.523080.40380 to 0.63978
gemma-4-26B-A4B-it0.3 to 0.431160.346560.516130.34840 to 0.68030
gemma-4-26B-A4B-it0.4 to 0.538200.441940.526320.37259 to 0.67521
gemma-4-26B-A4B-it0.5 to 0.631260.549150.838710.67366 to 0.92907
gemma-4-26B-A4B-it0.6 to 0.742350.650840.833330.69396 to 0.91684
gemma-4-26B-A4B-it0.7 to 0.835290.749330.828570.67318 to 0.91897
gemma-4-26B-A4B-it0.8 to 0.935310.835840.885710.74049 to 0.95465
gemma-4-26B-A4B-it0.9 to 1.0880.9249010.67559 to 1
jevify-gemma4-26b-a4b0.0 to 0.13,7733960.002928700.104960.09557 to 0.11514
jevify-gemma4-26b-a4b0.1 to 0.220170.124550.850000.63958 to 0.94763
jevify-gemma4-26b-a4b0.2 to 0.310100.2427910.72247 to 1
jevify-gemma4-26b-a4b0.3 to 0.4320.321020.666670.20766 to 0.93851
jevify-gemma4-26b-a4b0.4 to 0.5550.4303210.56552 to 1
jevify-gemma4-26b-a4b0.5 to 0.6110.5603110.20655 to 1
jevify-gemma4-26b-a4b0.6 to 0.7110.6775210.20655 to 1
jevify-gemma4-26b-a4b0.7 to 0.8110.7924510.20655 to 1
jevify-gemma4-26b-a4b0.9 to 1.0330.9539910.43850 to 1
open-jev-qwen-9b0.0 to 0.13,6793410.022600.092690.08374 to 0.10249
open-jev-qwen-9b0.1 to 0.259320.135120.542370.41658 to 0.66299
open-jev-qwen-9b0.2 to 0.317130.245630.764710.52738 to 0.90445
open-jev-qwen-9b0.3 to 0.41080.337850.800000.49016 to 0.94332
open-jev-qwen-9b0.4 to 0.5840.441960.500000.21522 to 0.78478
open-jev-qwen-9b0.5 to 0.6870.548780.875000.52911 to 0.97758
open-jev-qwen-9b0.6 to 0.7320.653670.666670.20766 to 0.93851
open-jev-qwen-9b0.7 to 0.811100.773060.909090.62264 to 0.98377
open-jev-qwen-9b0.8 to 0.9550.8486410.56552 to 1
open-jev-qwen-9b0.9 to 1.017140.954870.823530.58971 to 0.93809
DiffusionGemma0.0 to 0.13,5472630.008266980.074150.06598 to 0.08324
DiffusionGemma0.1 to 0.281420.140960.518520.41136 to 0.62400
DiffusionGemma0.2 to 0.352290.250210.557690.42340 to 0.68405
DiffusionGemma0.3 to 0.440280.341260.700000.54570 to 0.81925
DiffusionGemma0.4 to 0.521150.453790.714290.50044 to 0.86186
DiffusionGemma0.5 to 0.618130.539930.722220.49127 to 0.87500
DiffusionGemma0.6 to 0.71380.660250.615380.35523 to 0.82290
DiffusionGemma0.7 to 0.824180.753770.750000.55101 to 0.88001
DiffusionGemma0.8 to 0.9870.830240.875000.52911 to 0.97758
DiffusionGemma0.9 to 1.013130.9639210.77190 to 1
kev-9b0.0 to 0.13,3922640.040140.077830.06928 to 0.08733
kev-9b0.1 to 0.23711300.128930.350400.30361 to 0.40026
kev-9b0.2 to 0.344360.228380.818180.68039 to 0.90487
kev-9b0.3 to 0.4410.336430.250000.04559 to 0.69936
kev-9b0.4 to 0.5110.4008010.20655 to 1
kev-9b0.5 to 0.6210.576870.500000.09453 to 0.90547
kev-9b0.6 to 0.7220.6620410.34238 to 1
kev-9b0.9 to 1.0110.9516310.20655 to 1
open-jev-qwen-2b0.0 to 0.13,2733370.037090.102960.09301 to 0.11385
open-jev-qwen-2b0.1 to 0.2208310.138220.149040.10703 to 0.20378
open-jev-qwen-2b0.2 to 0.35970.245810.118640.05868 to 0.22524
open-jev-qwen-2b0.3 to 0.44030.350530.075000.02584 to 0.19864
open-jev-qwen-2b0.4 to 0.54030.443730.075000.02584 to 0.19864
open-jev-qwen-2b0.5 to 0.64770.542620.148940.07407 to 0.27686
open-jev-qwen-2b0.6 to 0.74580.650600.177780.09294 to 0.31330
open-jev-qwen-2b0.7 to 0.83580.750570.228570.12066 to 0.39017
open-jev-qwen-2b0.8 to 0.928100.855110.357140.20706 to 0.54170
open-jev-qwen-2b0.9 to 1.042220.947610.523810.37722 to 0.66640
decider-2b0.0 to 0.13,2672510.047460.076830.06819 to 0.08647
decider-2b0.1 to 0.23791110.133460.292880.24932 to 0.34059
decider-2b0.2 to 0.372270.242310.375000.27219 to 0.49047
decider-2b0.3 to 0.430160.343780.533330.36142 to 0.69768
decider-2b0.4 to 0.51360.440990.461540.23206 to 0.70856
decider-2b0.5 to 0.61860.553960.333330.16279 to 0.56251
decider-2b0.6 to 0.71680.655000.500000.28000 to 0.72000
decider-2b0.7 to 0.81150.731790.454550.21271 to 0.71991
decider-2b0.8 to 0.91160.857900.545450.28009 to 0.78729
bespoke-nimble-9b0.0 to 0.1900.087580.000000000.00000000 to 0.29915
bespoke-nimble-9b0.1 to 0.22,237210.157810.009387570.00614827 to 0.01431
bespoke-nimble-9b0.2 to 0.31,0741500.238170.139660.12022 to 0.16168
bespoke-nimble-9b0.3 to 0.43251830.346400.563080.50873 to 0.61595
bespoke-nimble-9b0.4 to 0.5118680.441060.576270.48609 to 0.66164
bespoke-nimble-9b0.5 to 0.638120.538190.315790.19085 to 0.47456
bespoke-nimble-9b0.6 to 0.71220.638140.166670.04697 to 0.44803
bespoke-nimble-9b0.7 to 0.8400.732690.000000000.00000000 to 0.48989
SecJudge0.0 to 0.1200.023250.000000000.00000000 to 0.65762
SecJudge0.2 to 0.34710.275450.021280.00376577 to 0.11113
SecJudge0.3 to 0.44000.309170.000000000.00000000 to 0.08762
SecJudge0.4 to 0.5100.485490.000000000.00000000 to 0.79345
SecJudge0.6 to 0.7515190.682520.036890.02374 to 0.05690
SecJudge0.7 to 0.839170.711970.017900.00869859 to 0.03649
SecJudge0.8 to 0.91,1921640.875640.137580.11919 to 0.15831
SecJudge0.9 to 1.01,6292450.934020.150400.13386 to 0.16858
Most of the corpus sits in the lowest bucket on every model, so the upper buckets are small and their intervals are wide. That is why the count is printed rather than left to the bar height. Source: outputs/curves/curves-s2.json :: arms.<arm>.calibration.buckets; except the log height of each count bar, which is a scale choice and is stated in the plot.

Notes on individual models

What the ranking rests on

Each row is a limit on how far the figures above travel.
AssumptionWhat it rests onWhat breaks it
The budget 0.00384502 is the right false-block price.It is OpenJev’s own realised block false-positive rate on this corpus, so it is a shipped operating point rather than a round number. At 3,381 benign cases it allows 13 false blocks.A deployment that tolerates more false blocks reorders the table: the ranking is a ranking at one budget, and the precision-recall curves show each model at every other budget it could be run at.
open-jev-qwen-27b is the best model.F1 0.73876 at the budget, against the next model’s 0.71162, with a bootstrap 95% interval on F1 of 0.70058 to 0.77309.The two intervals overlap, so this is a lead and not a separation. A rerun on a different case sample can reorder the top two.
Recall on this corpus means recall in deployment.382 of the 436 positives come from one source dataset, lihaonan0716/mcphunt-agent-traces, which is 87.61% of them.A deployment whose attack mix does not look like that source is not described by these recall figures. The per-source matrix on Method prints each model against each source with the positive count behind every rate.
A model’s F1 at the budget is a property of the model.Every figure at the budget is read on one variable, P(block) under aggregation definition A, at one threshold rule, so calibration is held fixed across the table.The shipped column mixes each model with its own calibration, which is why it is a column and not the ranking.
Accuracy is a useful summary here.It is not, at this prevalence: deciding allow on every case scores 0.88577. The accuracy figure is published beside that baseline and as a signed case count.Any accuracy claim that does not carry the all-allow baseline is uninformative on a corpus that is 88.58% benign.
The curves and the AUC column agree.12 of the 12 AUCs computed for the curves were compared against the figure the published re-mining pass recorded for the same arm, variable and definition, at a tolerance of 5e-15; 0 disagree.A rescore that moved an AUC without moving its curve would abort the pass that writes them, so the two cannot drift apart silently.

evaluation-only never-train
Nothing derived here is approved for training, synthetic generation, teacher context, distillation or redistribution. The 13 source datasets are public and linked on reproduce; the analysis artifacts behind each number are held privately and are available on request.