Deploy
How to run the cascade once the middle tier is chosen: which routing rule, which threshold, in what order, on what traffic, enforced at what granularity, and what it costs. Everything on this page uses OpenJev as the small model unless a panel names another.
Recommended: rules → OpenJev → Gemma 4
judge, two-sided routingThe small model settles a case alone when it is confident either way; only the uncertain band between the allow and block thresholds goes to the judge. at 0.30, with escalate-on-confirmWhen the rule engine answers confirm, that answer becomes a floor and the later tiers still run, so a case the small model or the judge would block ends as a block.: block-only F1
0.75173 at block FPR 0.00414, judge called on
15.95% of cases. The cascade as it runs today, which short-circuits on an
advisory rule confirm, scores 0.73773. The routing, threshold, cost and tier-order panels
below are measured on that version, because the escalating version was scored only in the
policy re-analysis. OpenJev
weights are CC BY-NC 4.0: non-commercial, attribution required, contact the authors for
commercial use. At production traffic, drop the judge (below).
The cascade
Three tiers, cheapest first. The rules settle only 13 of 3,817 cases, the small model settles 3,195, and 15.95% reach the judge.
Cascade pass-through rates
Data table
| Arrow / box | value |
|---|---|
| deterministic tier non-allow | 13 / 3,817 |
| LLM-judge call rate, one-sided routing | 0.90307 |
| LLM-judge call rate, two-sided @0.30 | 0.15955 |
| cascade block F1 (real deterministic tier) | 0.73773 |
| cascade block FPR | 0.00414 |
outputs/s2/policy-reanalysis-realdet.json :: compositions.realdet_short_circuit; outputs/deterministic-real/realdet-s2-openjev.json :: candidates[0].Broad comparison: four metrics per policy
Broad comparison: 4,277 scenarios, 3,817 scorable. OpenJev is the small model, Gemma 4 the judge. Four metrics, one panel each.
Data table
| Composition | block F1 | block FPR | confirm rate | LLM call rate |
|---|---|---|---|---|
| rules → Gemma 4 no System One tier | 0.71248 | 0.00680 | 0.43070 | 1.00000 |
| rules → OpenJev no LLM tier | 0.70231 | 0.00385 | 0.08488 | n/a — no LLM tier |
| rules → OpenJev → Gemma 4 one-sided routing | 0.73157 | 0.00592 | 0.43018 | 0.90307 |
| rules → OpenJev → Gemma 4 two-sided @0.30 | 0.73773 | 0.00414 | 0.16767 | 0.15955 |
| same two-sided cascade with escalate-on-confirm | 0.75173 | — | — | — |
rules → OpenJev ends 8.49% of cases as confirm with no judge to call; one-sided routing sends 90.31% of cases to the judge and still ends 43.02% as confirm. The last table row is the recommended escalate-on-confirm composition; its block-only F1 is read from a scorecard with an all-allow rule tier, which the policy re-analysis shows is identical to escalate-on-confirm on the block lens. Source: outputs/deterministic-real/realdet-s2-openjev.json :: candidates[0]; the last table row, outputs/s2/score-openjev.json; except the LLM-call rate for rules → Gemma 4 is 1.00000 by construction, because that composition has no System One tier to route past; the scorecard records no rate for it.Routing
one-sided routingThe small model settles only its confident blocks; every other case goes to the judge. trusts the small model only on its confident blocks. two-sided routingThe small model settles a case alone when it is confident either way; only the uncertain band between the allow and block thresholds goes to the judge. also trusts its confident allows, and that is the whole saving.
| Routing | block-only F1 | cases sent to the judge | confirm rate |
|---|---|---|---|
| one-sided | 0.73157 | 90.31% | 43.02% |
| two-sided @0.30 | 0.73773 | 15.95% | 16.77% |
The allow threshold
Four settings were run: 0.05, 0.10, 0.20, 0.30. The slider snaps to them; nothing in between was measured. For OpenJev the cheapest setting is also the best: block-only F1 moves +0.00616 across the four while the judge-call rate falls 32.04pp. That optimum is OpenJev’s: Jev 1.13.0 peaks at 0.20 (0.66366) and calls the judge 1.57× as often at 0.30.
Allow threshold against four metrics
Data table
| allow threshold | block F1 | block FPR | review rate | LLM-call rate | block F1, production-weighted |
|---|---|---|---|---|---|
| 0.05 | 0.73157 | 0.00592 | 0.3909 | 0.4800 | 0.06209 |
| 0.10 | 0.73259 | 0.00562 | 0.3395 | 0.3940 | 0.06385 |
| 0.20 | 0.73566 | 0.00473 | 0.2395 | 0.2544 | 0.07235 |
| 0.30 | 0.73773 | 0.00414 | 0.1677 | 0.1595 | 0.08215 |
outputs/deterministic-real/realdet-s2-openjev.json :: candidates[0].deterministic_then_system_one_then_llm_two_sided_*; outputs/deterministic-real/realdet-s3-openjev.json :: same.
Separate thresholds per surface save judge calls and nothing else: on one-shot
action calls a threshold of 0.5 cuts judge calls from
2.4658% to 1.37% of
1,482 cases at unchanged block-only F1 (0.80000);
stateful calls are already at their optimum at 0.30.
Tier order
| Cascade | block-only F1 | block FPR | cases sent to the judge |
|---|---|---|---|
| rules → OpenJev → judge | 0.73773 | 0.00414 | 15.95% |
| rules → OpenJev → DiffusionGemma → judge | 0.68956 | 0.00651 | 1.73% |
| rules → DiffusionGemma → OpenJev → judge | 0.48553 | 0.01035 | 1.73% |
| rules → DiffusionGemma → judge | 0.45574 | 0.01035 | 3.51% |
The stronger model goes first. A second small model in the right order gives up 0.04817 F1 for a further 9.2× cut in judge calls; in the wrong order it costs 0.20403.
Traffic mix: when to drop the judge
The Broad comparison is 88.58% benign. Real coding traffic is closer to the Production-weighted corpus, 99.10% benign. At that mix precision dominates F1, and a judge that escalates uncertain cases adds more false blocks than true ones.
Production-weighted: four metrics per policy
Production-weighted: 24,476 scenarios at the benign share of real traffic. Four metrics, one panel each.
Data table
| Composition | block F1 | block FPR | confirm rate | LLM call rate |
|---|---|---|---|---|
| rules → Gemma 4 no System One tier | 0.10980 | 0.09025 | 0.47422 | 1.00000 |
| rules → OpenJev no LLM tier | 0.21285 | 0.00924 | 0.06562 | n/a — no LLM tier |
| rules → OpenJev → Gemma 4 one-sided routing | 0.06134 | 0.07038 | 0.49718 | 0.95845 |
| rules → OpenJev → Gemma 4 two-sided @0.30 | 0.08215 | 0.04671 | 0.09695 | 0.10623 |
| same two-sided cascade with escalate-on-confirm | 0.12854 | — | — | — |
rules → OpenJev ends 6.56% of cases as confirm with no judge to call; one-sided routing sends 95.84% of cases to the judge and still ends 49.72% as confirm. The last table row is the recommended escalate-on-confirm composition; its block-only F1 is read from a scorecard with an all-allow rule tier, which the policy re-analysis shows is identical to escalate-on-confirm on the block lens. Source: outputs/deterministic-real/realdet-s3-openjev.json :: candidates[0]; the last table row, outputs/s3/score-openjev.json; except the LLM-call rate for rules → Gemma 4 is 1.00000 by construction, because that composition has no System One tier to route past; the scorecard records no rate for it.Production-weighted, block-only F1: full cascade 0.08215,
judge alone 0.10980, rules → OpenJev with no judge
0.21285 at block FPR 0.00924, a 6.56%
confirm rate and no judge calls.
With no judge, 3 small models were measured: Jev 1.13.0 at Q2, 0.23301 at block FPR 0.00590; OpenJev at Q2, 0.21285 at block FPR 0.00924; Jev 1.13.0 at Q3, 0.17729 at block FPR 0.00445; DiffusionGemma at Q2, 0.13068 at block FPR 0.08761; DiffusionGemma at Q3, 0.10717 at block FPR 0.05351. Jev 1.13.0 scores highest there, but it flags 0.33117 of real benign coding events against OpenJev’s 0.01944 and blocks 0.3111 of proof-verified compromises against OpenJev’s 0.5778. Dropping the judge also removes the only tier that fails closed (Risks).
Enforce per tool call, not per session
Per-call versus per-session enforcement
Same model, same threshold, same detection. The only change is whether one flagged call kills the session.
Data table
| Stage | model | per-event benign FPR | per-session benign FPR | ratio | unsafe recall | benign events blocked | benign sessions killed |
|---|---|---|---|---|---|---|---|
| Broad comparison | OpenJev | 0.000760 | 0.003845 | 5.06× | 0.55734 | 16 | 13 |
| Broad comparison | DiffusionGemma | 0.002708 | 0.010056 | 3.71× | 0.25917 | 57 | 34 |
| Production-weighted | OpenJev | 0.003144 | 0.009236 | 2.94× | 0.23982 | 313 | 224 |
| Production-weighted | DiffusionGemma | 0.019810 | 0.053515 | 2.70× | 0.38914 | 1,973 | 1,298 |
outputs/s2/policy-reanalysis-realdet.json, outputs/s3/policy-reanalysis-realdet.json :: compositions.realdet_short_circuit.per_event_vs_per_trajectory.Same model, same threshold. Stopping the one flagged call instead of the whole session costs 2.7× to 5.1× less benign work across the four measured model-and-corpus pairs: on the Broad comparison, 16 benign tool calls blocked against 13 whole benign sessions killed.
Cost against quality
The judge is the only metered tier: Gemma 4 on Bedrock costs $0.194 per 1,000 cases judged on the Broad comparison ($0.0923 on the Production-weighted corpus, whose prompts are shorter). Both small models ran self-hosted and recorded no provider spend. Each panel below compares policies within one model; the three models ran different question formats, named on each series.
Judge spend against block F1
One point per policy: six per small model per corpus, plus the judge alone, drawn once. Gemma 4 is the judge throughout.
Data table
| Corpus | Small model | Policy | judge spend per 1,000 cases | block-only F1 | block FPR | standing |
|---|---|---|---|---|---|---|
| Broad comparison | OpenJev (Q2) | rules → small model — no judge | $0.00000 (no judge) | 0.70231 | 0.00385 | undominated |
| Broad comparison | DiffusionGemma (Q3) | rules → small model — no judge | $0.00000 (no judge) | 0.38765 | 0.01006 | matched or beaten on both, with at least one strict loss |
| Broad comparison | Jev 1.13.0 (Q2) | rules → small model — no judge | $0.00000 (no judge) | 0.54153 | 0.00089 | matched or beaten on both, with at least one strict loss |
| Broad comparison | no small model | rules → judge — no small model | $0.194 | 0.71248 | 0.00680 | matched or beaten on both, with at least one strict loss |
| Broad comparison | OpenJev (Q2) | rules → small → judge — one-sided | $0.175 | 0.73157 | 0.00592 | matched or beaten on both, with at least one strict loss |
| Broad comparison | OpenJev (Q2) | rules → small → judge — two-sided @0.05 | $0.0929 | 0.73157 | 0.00592 | matched or beaten on both, with at least one strict loss |
| Broad comparison | OpenJev (Q2) | rules → small → judge — two-sided @0.10 | $0.0762 | 0.73259 | 0.00562 | matched or beaten on both, with at least one strict loss |
| Broad comparison | OpenJev (Q2) | rules → small → judge — two-sided @0.20 | $0.0492 | 0.73566 | 0.00473 | matched or beaten on both, with at least one strict loss |
| Broad comparison | OpenJev (Q2) | rules → small → judge — two-sided @0.30 | $0.0309 | 0.73773 | 0.00414 | undominated |
| Broad comparison | DiffusionGemma (Q3) | rules → small → judge — one-sided | $0.177 | 0.48896 | 0.01272 | matched or beaten on both, with at least one strict loss |
| Broad comparison | DiffusionGemma (Q3) | rules → small → judge — two-sided @0.05 | $0.0229 | 0.48882 | 0.01094 | matched or beaten on both, with at least one strict loss |
| Broad comparison | DiffusionGemma (Q3) | rules → small → judge — two-sided @0.10 | $0.0147 | 0.48475 | 0.01065 | matched or beaten on both, with at least one strict loss |
| Broad comparison | DiffusionGemma (Q3) | rules → small → judge — two-sided @0.20 | $0.00963 | 0.46753 | 0.01065 | matched or beaten on both, with at least one strict loss |
| Broad comparison | DiffusionGemma (Q3) | rules → small → judge — two-sided @0.30 | $0.00679 | 0.45574 | 0.01035 | matched or beaten on both, with at least one strict loss |
| Broad comparison | Jev 1.13.0 (Q2) | rules → small → judge — one-sided | $0.175 | 0.65970 | 0.00385 | matched or beaten on both, with at least one strict loss |
| Broad comparison | Jev 1.13.0 (Q2) | rules → small → judge — two-sided @0.05 | $0.117 | 0.65970 | 0.00385 | matched or beaten on both, with at least one strict loss |
| Broad comparison | Jev 1.13.0 (Q2) | rules → small → judge — two-sided @0.10 | $0.0992 | 0.66168 | 0.00325 | matched or beaten on both, with at least one strict loss |
| Broad comparison | Jev 1.13.0 (Q2) | rules → small → judge — two-sided @0.20 | $0.0660 | 0.66366 | 0.00266 | matched or beaten on both, with at least one strict loss |
| Broad comparison | Jev 1.13.0 (Q2) | rules → small → judge — two-sided @0.30 | $0.0486 | 0.65964 | 0.00266 | matched or beaten on both, with at least one strict loss |
| Production-weighted | OpenJev (Q2) | rules → small model — no judge | $0.00000 (no judge) | 0.21285 | 0.00924 | matched or beaten on both, with at least one strict loss |
| Production-weighted | DiffusionGemma (Q3) | rules → small model — no judge | $0.00000 (no judge) | 0.10717 | 0.05351 | matched or beaten on both, with at least one strict loss |
| Production-weighted | Jev 1.13.0 (Q2) | rules → small model — no judge | $0.00000 (no judge) | 0.23301 | 0.00590 | undominated |
| Production-weighted | no small model | rules → judge — no small model | $0.0923 | 0.10980 | 0.09025 | matched or beaten on both, with at least one strict loss |
| Production-weighted | OpenJev (Q2) | rules → small → judge — one-sided | $0.0884 | 0.06134 | 0.07038 | matched or beaten on both, with at least one strict loss |
| Production-weighted | OpenJev (Q2) | rules → small → judge — two-sided @0.05 | $0.0417 | 0.06209 | 0.06939 | matched or beaten on both, with at least one strict loss |
| Production-weighted | OpenJev (Q2) | rules → small → judge — two-sided @0.10 | $0.0256 | 0.06385 | 0.06465 | matched or beaten on both, with at least one strict loss |
| Production-weighted | OpenJev (Q2) | rules → small → judge — two-sided @0.20 | $0.0147 | 0.07235 | 0.05570 | matched or beaten on both, with at least one strict loss |
| Production-weighted | OpenJev (Q2) | rules → small → judge — two-sided @0.30 | $0.00980 | 0.08215 | 0.04671 | matched or beaten on both, with at least one strict loss |
| Production-weighted | DiffusionGemma (Q3) | rules → small → judge — one-sided | $0.0790 | 0.06445 | 0.07260 | matched or beaten on both, with at least one strict loss |
| Production-weighted | DiffusionGemma (Q3) | rules → small → judge — two-sided @0.05 | $0.0155 | 0.06810 | 0.06691 | matched or beaten on both, with at least one strict loss |
| Production-weighted | DiffusionGemma (Q3) | rules → small → judge — two-sided @0.10 | $0.0120 | 0.06713 | 0.06568 | matched or beaten on both, with at least one strict loss |
| Production-weighted | DiffusionGemma (Q3) | rules → small → judge — two-sided @0.20 | $0.00856 | 0.06783 | 0.06370 | matched or beaten on both, with at least one strict loss |
| Production-weighted | DiffusionGemma (Q3) | rules → small → judge — two-sided @0.30 | $0.00650 | 0.06470 | 0.06242 | matched or beaten on both, with at least one strict loss |
| Production-weighted | Jev 1.13.0 (Q2) | rules → small → judge — one-sided | $0.0838 | 0.07708 | 0.04745 | matched or beaten on both, with at least one strict loss |
| Production-weighted | Jev 1.13.0 (Q2) | rules → small → judge — two-sided @0.05 | $0.0557 | 0.07736 | 0.04725 | matched or beaten on both, with at least one strict loss |
| Production-weighted | Jev 1.13.0 (Q2) | rules → small → judge — two-sided @0.10 | $0.0429 | 0.07835 | 0.04651 | matched or beaten on both, with at least one strict loss |
| Production-weighted | Jev 1.13.0 (Q2) | rules → small → judge — two-sided @0.20 | $0.0304 | 0.08227 | 0.04374 | matched or beaten on both, with at least one strict loss |
| Production-weighted | Jev 1.13.0 (Q2) | rules → small → judge — two-sided @0.30 | $0.0226 | 0.08682 | 0.04086 | matched or beaten on both, with at least one strict loss |
outputs/deterministic-real/realdet-s2-openjev.json, outputs/deterministic-real/realdet-s3-openjev.json :: candidates[0].<policy>.{binary_block_only.f1, llm_invocation_rate}; judge price from outputs/s2/gemma4-c7.jsonl.meta.json :: prompt_tokens / cases / provider_calls (the Q0 judge run the cascade was scored from, which records no price) at the rate in outputs/s2/gemma4-q2.jsonl.meta.json :: estimated_usd / prompt_tokens (the Q2 judge run, the only one priced) and outputs/s3/gemma4-c7.jsonl.meta.json :: prompt_tokens / cases / provider_calls (the Q0 judge run the cascade was scored from, which records no price) at the rate in outputs/s2/gemma4-q2.jsonl.meta.json :: estimated_usd / prompt_tokens (the Q2 judge run, the only one priced); outputs/deterministic-real/realdet-s2-jev.json :: candidates[0] (jev-1.13.0/C7/I3/Q2); outputs/deterministic-real/realdet-s3-jev.json :: candidates[0] (jev-1.13.0/C7/I3/Q2).The cascade as a flow
Data table
| Junction | cases | share of all cases |
|---|---|---|
| all cases entering | 3,817 | 1 |
| terminated by the deterministic rules | 13 | 0.003406 |
| passed through the rules to the small model | 3,804 | 0.996594 |
| decided by the small model, judge never called | 3,195 | 0.837045 |
| escalated to the judge | 609 | 0.159549 |
| ends allow | 2,900 | 0.759759 |
| ends confirm | 640 | 0.167671 |
| ends block | 277 | 0.072570 |
outputs/s2/policy-reanalysis-realdet.json :: compositions/realdet_short_circuit/cascade_tiers/two_tier_openjev_then_gemma.{decided_by, gemma_invocation_rate, counts}; final mix from outputs/deterministic-real/realdet-s2-openjev.json :: candidates[0].deterministic_then_system_one_then_llm_two_sided_0.30.three_way.confusion.Review burden against attacks allowed through
Data table
| Model | allow threshold | confirm rate | unsafe allowed through | block-only F1 |
|---|---|---|---|---|
| OpenJev | 0.05 | 0.390883 | 0.087156 | 0.73157 |
| OpenJev | 0.10 | 0.339534 | 0.094037 | 0.73259 |
| OpenJev | 0.20 | 0.239455 | 0.116972 | 0.73566 |
| OpenJev | 0.30 | 0.167671 | 0.139908 | 0.73773 |
| DiffusionGemma | 0.05 | 0.148022 | 0.165138 | 0.48882 |
| DiffusionGemma | 0.10 | 0.112392 | 0.183486 | 0.48475 |
| DiffusionGemma | 0.20 | 0.088289 | 0.206422 | 0.46753 |
| DiffusionGemma | 0.30 | 0.075452 | 0.222477 | 0.45574 |
| Jev 1.13.0 | 0.05 | 0.439874 | 0.094037 | 0.65970 |
| Jev 1.13.0 | 0.10 | 0.423369 | 0.114679 | 0.66168 |
| Jev 1.13.0 | 0.20 | 0.330364 | 0.126147 | 0.66366 |
| Jev 1.13.0 | 0.30 | 0.269059 | 0.146789 | 0.65964 |
| rules → judge (reference) | — | 0.430705 | 0.094037 | 0.71248 |
outputs/deterministic-real/realdet-s2-openjev.json, outputs/deterministic-real/realdet-s2-diffgemma.json :: candidates[0].deterministic_then_system_one_then_llm_two_sided_* :: {review_rate, binary.confusion}; outputs/deterministic-real/realdet-s2-jev.json :: candidates[0] (jev-1.13.0/C7/I3/Q2).Price it against your own traffic
The projection prices the judge only, from the measured rates: not your own tools, retries, in-line latency or the human cost of a confirm.
| Interpolated rate | value |
|---|---|
| share of cases that reach the judge | 0.106227 |
| share of cases ending confirm | 0.096952 |
| share of unsafe cases ending allow | 0.045249 |
| cases per day, as entered | 14,000 |
| benign share actually used (clamped to the measured range) | 0.9910 |
Serving cost and latency
Recorded provider spend over every run on this site is $23.95: $23.18 for hosted Jev 1.13.0 (551,834,284 input tokens at $0.042 per million, the one rate every priced manifest records) and $0.777 for the Gemma 4 judge, whose large-stage runs logged no cost, so its true Bedrock spend is higher by an unknown amount. The self-hosted models ran on one g6e.12xlarge (published on-demand rate $12.90 per hour); its uptime is not recorded, so no GPU total is stated.
The small model emits one token per call, so serving time is all prefill: there is no decode phase to speed up and input tokens are the whole cost. The latencies below come from a saturated GPU running a batch to completion; the long tail is queueing, and any in-line deployment needs a timeout policy.
Wall-clock latency under batch load
Wall-clock times per case. The self-hosted rows come from a saturated GPU running a batch to completion, so a single in-line call would see less; a hosted-API row is a network round trip and is not comparable with them.
Data table
| Run | decisions | p50 (s) | p95 (s) | p99 (s) | max (s) | spend | input tokens | errors |
|---|---|---|---|---|---|---|---|---|
| Broad comparison · OpenJev | 3,817 | 19.44 | 120.52 | 252.70 | 434.1 | $1.97 | 46,902,085 | 0 |
| Broad comparison · DiffusionGemma | 3,817 | 3.66 | 23.87 | 47.99 | 93.6 | $1.29 | 30,791,741 | 0 |
| Broad comparison · Jev 1.13.0 | 3,817 | 1.42 | 7.35 | 14.45 | 21.2 | $1.21 | 28,859,889 | 0 |
| Production-weighted · OpenJev | 24,476 | 5.19 | 42.23 | 47.75 | 365.4 | $5.34 | 127,247,454 | 32 |
| Production-weighted · DiffusionGemma | 24,476 | 1.36 | 10.63 | 13.60 | 75.4 | $4.06 | 96,676,729 | 0 |
| Production-weighted · Jev 1.13.0 | 24,476 | 0.40 | 3.07 | 3.42 | 22.0 | $3.72 | 88,682,752 | 0 |
outputs/deterministic-real/realdet-s2-openjev.json; outputs/deterministic-real/realdet-s2-diffgemma.json; outputs/deterministic-real/realdet-s2-jev.json; outputs/deterministic-real/realdet-s3-openjev.json; outputs/deterministic-real/realdet-s3-diffgemma.json; outputs/deterministic-real/realdet-s3-jev.json :: candidates[0].system_one.latency_ms / estimated_usd / input_tokens / errors.| Metric | default | flex | change |
|---|---|---|---|
| median per-case latency | 844.6 ms | 813.8 ms | −3.6% |
| p95 per-case latency | 23,502.7 ms | 15,908.9 ms | −32.3% |
| worst case | 122,211 ms | 33,560 ms | −72.5% |
| tokens, errors | 993,812 / 3,038, 0 | 993,812 / 3,038, 0 | identical |
| case decisions | 199 of 200 identical | ||
Prefix caching does not pay for this workload. 82.14% of 1,008 measured prompts are shorter than one 784-token cache block. Padding prompts to the block boundary cost +86.68% input tokens. A finer cache match unit raised hit rates but slowed the server: on the 120-prompt probe by 34.7% serially and 15.6% at four concurrent requests, and on the 200-case benchmark workload by 0.72% of wall clock. Putting the static part of the prompt first costs nothing.
Recommendations
| # | Recommendation | Evidence | Confidence |
|---|---|---|---|
| 1 | Enforce per tool call, not per session. | 2.7×–5.1× less benign work destroyed at identical recall (above). | high |
| 2 | Use a single-disposition question (Q2, or Q4) for anything that blocks; never the boolean-probe formats Q1 and Q3. | Block-only F1 0.70231 against 0.03160 and 0.03604; Q1 and Q3 fail open on a malformed answer (Risks). | high |
| 3 | Set the escalation threshold per surface to save judge calls. | 1.0959 points fewer judge calls on one-shot actions at unchanged F1. | qualified |
| 4 | Treat a rule-engine confirm as a floor and keep evaluating. | 0.73773 → 0.75173 block-only F1, same false blocks and judge-call rate. | high |
| 5 | At production traffic, run rules → small model with no judge, once recommendation 6 is in place. | 0.08215 → 0.21285 block-only F1 at zero judge cost (above). | high |
| 6 | Make the disposition on a small-model failure configurable, defaulting to confirm, with a timeout. | 24 of 28 injected failure classes end with the call allowed (Risks). | high |
| 7 | Decide explicitly whether to accept the short-circuit loss of two-sided routing. | Of the 1,346 cases the judge wanted to escalate and the small model allowed, 32 are graded unsafe (2.38%); two-sided routing discards at most that set (Risks). | medium |
| 8 | Author the mined deterministic rules, starting with shell-history tampering. | 96.2% of the small model’s confident blocks match no rule; that cluster is grade A throughout (20 events). | medium |
| 9 | Fix the two provenance defects in the dataset lock. | One disabled source was normalised into the intent corpus; one entry pins a paper, not a revision (Method). | high |
| 10 | Do not use the serves-the-request score (Lane B) as a standalone gate. | It flags 0.17887 of benign tool calls and 39/40 benign coding sessions (Risks). | high |
| 11 | Do not invest in prefix caching or prompt padding. | +86.68% input tokens for padding; slower serving with a finer match unit (above). | qualified |
| 12 | Do not reinstate Von 1.0.1. | Block-only F1 0.76786 against 0.71545 for blocking everything on the same pilot, at a 38.8s–44.4s median per case on CPU. | high |
What would change these numbers
| Assumption | What breaks it |
|---|---|
| The judge price, $0.194 per 1,000 cases judged | A different context length, service tier or provider price. |
| The benign share, between 88.58% and 99.10% | A third mix; nothing between or outside the two was run. |
A confirm stops nothing | If your confirm path is enforced by a reliable second decider, any-intervention F1F1 counting a confirm as a catch as well as a block. It measures triage; it overstates what a guard actually stops. becomes the relevant score and the order of policies changes. |
| Recall on proven compromises | Even the best model blocks only 0.5778 of proof-verified compromises (Intent). No setting here makes the guard sufficient on its own. |
evaluation-only never-train
Nothing derived here is approved for training, synthetic generation, teacher context, distillation or redistribution. The 13 source datasets are public and linked on reproduce; the analysis artifacts behind each number are held privately and are available on request.