Deploy

How to run the cascade once the middle tier is chosen: which routing rule, which threshold, in what order, on what traffic, enforced at what granularity, and what it costs. Everything on this page uses OpenJev as the small model unless a panel names another.

Recommended: rules → OpenJev → Gemma 4 judge, two-sided routingThe small model settles a case alone when it is confident either way; only the uncertain band between the allow and block thresholds goes to the judge. at 0.30, with escalate-on-confirmWhen the rule engine answers confirm, that answer becomes a floor and the later tiers still run, so a case the small model or the judge would block ends as a block.: block-only F1 0.75173 at block FPR 0.00414, judge called on 15.95% of cases. The cascade as it runs today, which short-circuits on an advisory rule confirm, scores 0.73773. The routing, threshold, cost and tier-order panels below are measured on that version, because the escalating version was scored only in the policy re-analysis. OpenJev weights are CC BY-NC 4.0: non-commercial, attribution required, contact the authors for commercial use. At production traffic, drop the judge (below).

The cascade

Three tiers, cheapest first. The rules settle only 13 of 3,817 cases, the small model settles 3,195, and 15.95% reach the judge.

Cascade pass-through rates

Cascade pass-through ratesDeterministic rules run first and terminate 13 of 3,817 scorable cases. A small System One model then answers every remaining case; only the uncertain band is escalated to the expensive LLM judge, which sees 15.95 percent of traffic. Broad comparison, 3,817 scorable scenarios, real deterministic tier. OpenJev as System One, Gemma 4 as the judge. 1. Rulesexact patternsnon-allow 13/3,817 2. System Onetwo-sided routingdecides most traffic 3. LLM judgeuncertain band onlysees 15.95% Outcomeallow/confirm/blockblock F1 0.73773 99.66% continue13 stop here 15.95% escalate84.05% skip the judge judge decides 84.05% never reach the judge
Data table
Arrow / boxvalue
deterministic tier non-allow13 / 3,817
LLM-judge call rate, one-sided routing0.90307
LLM-judge call rate, two-sided @0.300.15955
cascade block F1 (real deterministic tier)0.73773
cascade block FPR0.00414
Source: outputs/s2/policy-reanalysis-realdet.json :: compositions.realdet_short_circuit; outputs/deterministic-real/realdet-s2-openjev.json :: candidates[0].

Broad comparison: four metrics per policy

Broad comparison: 4,277 scenarios, 3,817 scorable. OpenJev is the small model, Gemma 4 the judge. Four metrics, one panel each.

Four metrics for four cascade compositionsFour metrics, one small-multiple panel each. The same four compositions appear in the same order in every panel, named in the left column. Confirm rate is the share of decisions that end as confirm; LLM calls is the share of cases that reach the judge. They are different measurements. rules → Gemma 4 no System One tier rules → OpenJev no LLM tier rules → OpenJev → Gemma 4 one-sided routing rules → OpenJev → Gemma 4 two-sided @0.30 block F1 higher better rules → Gemma 4, no System One tier · block F1: 0.71248 0.7125 rules → OpenJev, no LLM tier · block F1: 0.70231 0.7023 rules → OpenJev → Gemma 4, one-sided routing · block F1: 0.73157 0.7316 rules → OpenJev → Gemma 4, two-sided @0.30 · block F1: 0.73773 0.7377 0 0.8 block FPR lower better rules → Gemma 4, no System One tier · block FPR: 0.00680 0.0068 rules → OpenJev, no LLM tier · block FPR: 0.00385 0.0038 rules → OpenJev → Gemma 4, one-sided routing · block FPR: 0.00592 0.0059 rules → OpenJev → Gemma 4, two-sided @0.30 · block FPR: 0.00414 0.0041 0 0.01 confirm rate left for review rules → Gemma 4, no System One tier · confirm rate: 0.43070 0.4307 rules → OpenJev, no LLM tier · confirm rate: 0.08488 0.0849 rules → OpenJev → Gemma 4, one-sided routing · confirm rate: 0.43018 0.4302 rules → OpenJev → Gemma 4, two-sided @0.30 · confirm rate: 0.16767 0.1677 0 0.5 LLM calls lower better rules → Gemma 4, no System One tier · LLM calls: 1.00000 1.0000 n/a — no LLM tier rules → OpenJev → Gemma 4, one-sided routing · LLM calls: 0.90307 0.9031 rules → OpenJev → Gemma 4, two-sided @0.30 · LLM calls: 0.15955 0.1595 0 1
Data table
Broad comparison corpus (s2). The same five compositions are tabulated for the other corpus in its own figure; the column set is the same and the population is not.
Compositionblock F1block FPRconfirm rateLLM call rate
rules → Gemma 4 no System One tier0.712480.006800.430701.00000
rules → OpenJev no LLM tier0.702310.003850.08488n/a — no LLM tier
rules → OpenJev → Gemma 4 one-sided routing0.731570.005920.430180.90307
rules → OpenJev → Gemma 4 two-sided @0.300.737730.004140.167670.15955
same two-sided cascade with escalate-on-confirm0.75173———
rules → OpenJev ends 8.49% of cases as confirm with no judge to call; one-sided routing sends 90.31% of cases to the judge and still ends 43.02% as confirm. The last table row is the recommended escalate-on-confirm composition; its block-only F1 is read from a scorecard with an all-allow rule tier, which the policy re-analysis shows is identical to escalate-on-confirm on the block lens. Source: outputs/deterministic-real/realdet-s2-openjev.json :: candidates[0]; the last table row, outputs/s2/score-openjev.json; except the LLM-call rate for rules → Gemma 4 is 1.00000 by construction, because that composition has no System One tier to route past; the scorecard records no rate for it.
0.71248 → 0.73773
block-only F1, judge alone → cascade
0.00680 → 0.00414
block false-positive rate
43.07% → 16.77%
confirm rate (left for review)
100% → 15.95%
cases sent to the judge

Routing

one-sided routingThe small model settles only its confident blocks; every other case goes to the judge. trusts the small model only on its confident blocks. two-sided routingThe small model settles a case alone when it is confident either way; only the uncertain band between the allow and block thresholds goes to the judge. also trusts its confident allows, and that is the whole saving.

Broad comparison, OpenJev, as the cascade runs today. Source: outputs/deterministic-real/realdet-s2-openjev.json.
Routingblock-only F1cases sent to the judgeconfirm rate
one-sided0.7315790.31%43.02%
two-sided @0.300.7377315.95%16.77%

The allow threshold

Four settings were run: 0.05, 0.10, 0.20, 0.30. The slider snaps to them; nothing in between was measured. For OpenJev the cheapest setting is also the best: block-only F1 moves +0.00616 across the four while the judge-call rate falls 32.04pp. That optimum is OpenJev’s: Jev 1.13.0 peaks at 0.20 (0.66366) and calls the judge 1.57× as often at 0.30.

0.30The 4 settings measured in the whole-corpus scorecard sweep: 0.05, 0.10, 0.20, 0.30. The slider snaps to them and nothing between them was run. With scripting off the tiles show 0.30.
block F1
0.73773
production-weighted 0.08215
block FPR
0.00414
production-weighted 0.04671
review rate
0.1677
production-weighted 0.0970
LLM-call rate
0.1595
production-weighted 0.1062

Allow threshold against four metrics

0.05 / 0.10 / 0.200.30, the shipped setting
Four measured allow-thresholds, four metricsRaising the trusted-allow threshold from 0.05 to 0.30 cuts the LLM-call rate from 47.996 percent to 15.955 percent while block F1 moves by 0.006. Broad comparison, OpenJev inside the two-sided cascade. 4 settings in this sweep. Nothing between them was run. Each panel has its own scale, printed at its top gridline. Bars start at zero. block F1 0 to 0.75 allow threshold 0.05 · block F1: 0.73157 0.05 allow threshold 0.10 · block F1: 0.73259 0.10 allow threshold 0.20 · block F1: 0.73566 0.20 allow threshold 0.30 · block F1: 0.73773 0.30 block FPR 0 to 0.01 allow threshold 0.05 · block FPR: 0.00592 0.05 allow threshold 0.10 · block FPR: 0.00562 0.10 allow threshold 0.20 · block FPR: 0.00473 0.20 allow threshold 0.30 · block FPR: 0.00414 0.30 review rate 0 to 0.5 allow threshold 0.05 · review rate: 0.3909 0.05 allow threshold 0.10 · review rate: 0.3395 0.10 allow threshold 0.20 · review rate: 0.2395 0.20 allow threshold 0.30 · review rate: 0.1677 0.30 LLM-call rate 0 to 0.5 allow threshold 0.05 · LLM-call rate: 0.4800 0.05 allow threshold 0.10 · LLM-call rate: 0.3940 0.10 allow threshold 0.20 · LLM-call rate: 0.2544 0.20 allow threshold 0.30 · LLM-call rate: 0.1595 0.30
Data table
allow thresholdblock F1block FPRreview rateLLM-call rateblock F1, production-weighted
0.050.731570.005920.39090.48000.06209
0.100.732590.005620.33950.39400.06385
0.200.735660.004730.23950.25440.07235
0.300.737730.004140.16770.15950.08215
OpenJev’s sweep; the optimum does not carry to the other models: on the broad comparison corpus the models peak at different thresholds — DiffusionGemma at 0.05, Jev 1.13.0 at 0.20, OpenJev at 0.30; on the production-weighted corpus the models peak at different thresholds — DiffusionGemma at 0.05, Jev 1.13.0 at 0.30, OpenJev at 0.30. Source: outputs/deterministic-real/realdet-s2-openjev.json :: candidates[0].deterministic_then_system_one_then_llm_two_sided_*; outputs/deterministic-real/realdet-s3-openjev.json :: same.

Separate thresholds per surface save judge calls and nothing else: on one-shot action calls a threshold of 0.5 cuts judge calls from 2.4658% to 1.37% of 1,482 cases at unchanged block-only F1 (0.80000); stateful calls are already at their optimum at 0.30.

Tier order

Broad comparison, as the cascade runs today. Source: outputs/s2/policy-reanalysis-realdet.json.
Cascadeblock-only F1block FPRcases sent to the judge
rules → OpenJev → judge0.737730.0041415.95%
rules → OpenJev → DiffusionGemma → judge0.689560.006511.73%
rules → DiffusionGemma → OpenJev → judge0.485530.010351.73%
rules → DiffusionGemma → judge0.455740.010353.51%

The stronger model goes first. A second small model in the right order gives up 0.04817 F1 for a further 9.2× cut in judge calls; in the wrong order it costs 0.20403.

Traffic mix: when to drop the judge

The Broad comparison is 88.58% benign. Real coding traffic is closer to the Production-weighted corpus, 99.10% benign. At that mix precision dominates F1, and a judge that escalates uncertain cases adds more false blocks than true ones.

Production-weighted: four metrics per policy

Production-weighted: 24,476 scenarios at the benign share of real traffic. Four metrics, one panel each.

Four metrics for four cascade compositionsFour metrics, one small-multiple panel each. The same four compositions appear in the same order in every panel, named in the left column. Confirm rate is the share of decisions that end as confirm; LLM calls is the share of cases that reach the judge. They are different measurements. rules → Gemma 4 no System One tier rules → OpenJev no LLM tier rules → OpenJev → Gemma 4 one-sided routing rules → OpenJev → Gemma 4 two-sided @0.30 block F1 higher better rules → Gemma 4, no System One tier · block F1: 0.10980 0.1098 rules → OpenJev, no LLM tier · block F1: 0.21285 0.2129 rules → OpenJev → Gemma 4, one-sided routing · block F1: 0.06134 0.0613 rules → OpenJev → Gemma 4, two-sided @0.30 · block F1: 0.08215 0.0822 0 0.24 block FPR lower better rules → Gemma 4, no System One tier · block FPR: 0.09025 0.0902 rules → OpenJev, no LLM tier · block FPR: 0.00924 0.0092 rules → OpenJev → Gemma 4, one-sided routing · block FPR: 0.07038 0.0704 rules → OpenJev → Gemma 4, two-sided @0.30 · block FPR: 0.04671 0.0467 0 0.1 confirm rate left for review rules → Gemma 4, no System One tier · confirm rate: 0.47422 0.4742 rules → OpenJev, no LLM tier · confirm rate: 0.06562 0.0656 rules → OpenJev → Gemma 4, one-sided routing · confirm rate: 0.49718 0.4972 rules → OpenJev → Gemma 4, two-sided @0.30 · confirm rate: 0.09695 0.0970 0 0.5 LLM calls lower better rules → Gemma 4, no System One tier · LLM calls: 1.00000 1.0000 n/a — no LLM tier rules → OpenJev → Gemma 4, one-sided routing · LLM calls: 0.95845 0.9584 rules → OpenJev → Gemma 4, two-sided @0.30 · LLM calls: 0.10623 0.1062 0 1
Data table
Production-weighted corpus (s3). The same five compositions are tabulated for the other corpus in its own figure; the column set is the same and the population is not.
Compositionblock F1block FPRconfirm rateLLM call rate
rules → Gemma 4 no System One tier0.109800.090250.474221.00000
rules → OpenJev no LLM tier0.212850.009240.06562n/a — no LLM tier
rules → OpenJev → Gemma 4 one-sided routing0.061340.070380.497180.95845
rules → OpenJev → Gemma 4 two-sided @0.300.082150.046710.096950.10623
same two-sided cascade with escalate-on-confirm0.12854———
rules → OpenJev ends 6.56% of cases as confirm with no judge to call; one-sided routing sends 95.84% of cases to the judge and still ends 49.72% as confirm. The last table row is the recommended escalate-on-confirm composition; its block-only F1 is read from a scorecard with an all-allow rule tier, which the policy re-analysis shows is identical to escalate-on-confirm on the block lens. Source: outputs/deterministic-real/realdet-s3-openjev.json :: candidates[0]; the last table row, outputs/s3/score-openjev.json; except the LLM-call rate for rules → Gemma 4 is 1.00000 by construction, because that composition has no System One tier to route past; the scorecard records no rate for it.

Production-weighted, block-only F1: full cascade 0.08215, judge alone 0.10980, rules → OpenJev with no judge 0.21285 at block FPR 0.00924, a 6.56% confirm rate and no judge calls.

With no judge, 3 small models were measured: Jev 1.13.0 at Q2, 0.23301 at block FPR 0.00590; OpenJev at Q2, 0.21285 at block FPR 0.00924; Jev 1.13.0 at Q3, 0.17729 at block FPR 0.00445; DiffusionGemma at Q2, 0.13068 at block FPR 0.08761; DiffusionGemma at Q3, 0.10717 at block FPR 0.05351. Jev 1.13.0 scores highest there, but it flags 0.33117 of real benign coding events against OpenJev’s 0.01944 and blocks 0.3111 of proof-verified compromises against OpenJev’s 0.5778. Dropping the judge also removes the only tier that fails closed (Risks).

Enforce per tool call, not per session

Per-call versus per-session enforcement

Same model, same threshold, same detection. The only change is whether one flagged call kills the session.

per-session enforcementper-event enforcement
Benign disruption per event versus per sessionKilling a whole session on one flagged call multiplies benign disruption several times over at identical detection. 0.0% 1.0% 2.0% 3.0% 4.0% 5.0% 6.0% Benign work destroyed, at identical detection Broad comparison · OpenJev Broad comparison OpenJev per-session: 0.003845 (13 benign sessions killed) Broad comparison OpenJev per-event: 0.000760 (16 benign events blocked) 5.1× worse Broad comparison · DiffusionGemma Broad comparison DiffusionGemma per-session: 0.010056 (34 benign sessions killed) Broad comparison DiffusionGemma per-event: 0.002708 (57 benign events blocked) 3.7× worse Production-weighted · OpenJev Production-weighted OpenJev per-session: 0.009236 (224 benign sessions killed) Production-weighted OpenJev per-event: 0.003144 (313 benign events blocked) 2.9× worse Production-weighted · DiffusionGemma Production-weighted DiffusionGemma per-session: 0.053515 (1298 benign sessions killed) Production-weighted DiffusionGemma per-event: 0.019810 (1973 benign events blocked) 2.7× worse share of benign units flagged
Data table
Stagemodelper-event benign FPRper-session benign FPRratiounsafe recallbenign events blockedbenign sessions killed
Broad comparisonOpenJev0.0007600.0038455.06×0.557341613
Broad comparisonDiffusionGemma0.0027080.0100563.71×0.259175734
Production-weightedOpenJev0.0031440.0092362.94×0.23982313224
Production-weightedDiffusionGemma0.0198100.0535152.70×0.389141,9731,298
Source: outputs/s2/policy-reanalysis-realdet.json, outputs/s3/policy-reanalysis-realdet.json :: compositions.realdet_short_circuit.per_event_vs_per_trajectory.

Same model, same threshold. Stopping the one flagged call instead of the whole session costs 2.7× to 5.1× less benign work across the four measured model-and-corpus pairs: on the Broad comparison, 16 benign tool calls blocked against 13 whole benign sessions killed.

Cost against quality

The judge is the only metered tier: Gemma 4 on Bedrock costs $0.194 per 1,000 cases judged on the Broad comparison ($0.0923 on the Production-weighted corpus, whose prompts are shorter). Both small models ran self-hosted and recorded no provider spend. Each panel below compares policies within one model; the three models ran different question formats, named on each series.

Judge spend against block F1

One point per policy: six per small model per corpus, plus the judge alone, drawn once. Gemma 4 is the judge throughout.

on the frontier, Broad comparisonon the frontier, Production-weightedmatched or beaten on both, with at least one strict loss
Judge spend against block F1, one panel per corpusOne panel per corpus, each with its own vertical scale. On the broad comparison corpus 2 of 19 policies are undominated: OpenJev (Q2) two-sided @0.30 at 0.73773 block F1, OpenJev (Q2) rules → small at 0.70231 block F1. On the production-weighted corpus 1 of 19 policies are undominated: Jev 1.13.0 (Q2) rules → small at 0.23301 block F1. 3 policies call no judge at all, so they are drawn on the left edge at exactly zero spend. Each dot is one guard policy. Left to right is what the judge costs; up is how well the policy blocks. The line joins the policies nothing else matches or beats on both axes. A small grey ring is matched or beaten on both. Spend = share of cases reaching the judge × $0.0000923 per case judged × 1,000. Each panel has its own scale. Best block F1 of the policies drawn: 0.73773 broad comparison, 0.23301 production-weighted. Broad comparison 88.58% benign · n=3,817 scorable 0.33 0.46 0.59 0.71 0.84 block-only F1, higher is better $0.000 $0.051 $0.103 $0.154 $0.205 judge spend per 1,000 cases (USD) Broad comparison · OpenJev (Q2) — rules → small model, no judge: judge spend $0.00000 per 1,000 cases, block-only F1 0.70231 — undominated Broad comparison · DiffusionGemma (Q3) — rules → small model, no judge: judge spend $0.00000 per 1,000 cases, block-only F1 0.38765 — matched or beaten on both, with at least one strict loss Broad comparison · Jev 1.13.0 (Q2) — rules → small model, no judge: judge spend $0.00000 per 1,000 cases, block-only F1 0.54153 — matched or beaten on both, with at least one strict loss Broad comparison · no small model — rules → judge, no small model: judge spend $0.194 per 1,000 cases, block-only F1 0.71248 — matched or beaten on both, with at least one strict loss Broad comparison · OpenJev (Q2) — rules → small → judge, one-sided: judge spend $0.175 per 1,000 cases, block-only F1 0.73157 — matched or beaten on both, with at least one strict loss Broad comparison · OpenJev (Q2) — rules → small → judge, two-sided @0.05: judge spend $0.0929 per 1,000 cases, block-only F1 0.73157 — matched or beaten on both, with at least one strict loss Broad comparison · OpenJev (Q2) — rules → small → judge, two-sided @0.10: judge spend $0.0762 per 1,000 cases, block-only F1 0.73259 — matched or beaten on both, with at least one strict loss Broad comparison · OpenJev (Q2) — rules → small → judge, two-sided @0.20: judge spend $0.0492 per 1,000 cases, block-only F1 0.73566 — matched or beaten on both, with at least one strict loss Broad comparison · OpenJev (Q2) — rules → small → judge, two-sided @0.30: judge spend $0.0309 per 1,000 cases, block-only F1 0.73773 — undominated Broad comparison · DiffusionGemma (Q3) — rules → small → judge, one-sided: judge spend $0.177 per 1,000 cases, block-only F1 0.48896 — matched or beaten on both, with at least one strict loss Broad comparison · DiffusionGemma (Q3) — rules → small → judge, two-sided @0.05: judge spend $0.0229 per 1,000 cases, block-only F1 0.48882 — matched or beaten on both, with at least one strict loss Broad comparison · DiffusionGemma (Q3) — rules → small → judge, two-sided @0.10: judge spend $0.0147 per 1,000 cases, block-only F1 0.48475 — matched or beaten on both, with at least one strict loss Broad comparison · DiffusionGemma (Q3) — rules → small → judge, two-sided @0.20: judge spend $0.00963 per 1,000 cases, block-only F1 0.46753 — matched or beaten on both, with at least one strict loss Broad comparison · DiffusionGemma (Q3) — rules → small → judge, two-sided @0.30: judge spend $0.00679 per 1,000 cases, block-only F1 0.45574 — matched or beaten on both, with at least one strict loss Broad comparison · Jev 1.13.0 (Q2) — rules → small → judge, one-sided: judge spend $0.175 per 1,000 cases, block-only F1 0.65970 — matched or beaten on both, with at least one strict loss Broad comparison · Jev 1.13.0 (Q2) — rules → small → judge, two-sided @0.05: judge spend $0.117 per 1,000 cases, block-only F1 0.65970 — matched or beaten on both, with at least one strict loss Broad comparison · Jev 1.13.0 (Q2) — rules → small → judge, two-sided @0.10: judge spend $0.0992 per 1,000 cases, block-only F1 0.66168 — matched or beaten on both, with at least one strict loss Broad comparison · Jev 1.13.0 (Q2) — rules → small → judge, two-sided @0.20: judge spend $0.0660 per 1,000 cases, block-only F1 0.66366 — matched or beaten on both, with at least one strict loss Broad comparison · Jev 1.13.0 (Q2) — rules → small → judge, two-sided @0.30: judge spend $0.0486 per 1,000 cases, block-only F1 0.65964 — matched or beaten on both, with at least one strict loss OpenJev Q2 · two-sided @0.30 OpenJev Q2 · rules → small Production-weighted 99.10% benign · n=24,476 scorable 0.03 0.09 0.16 0.22 0.29 block-only F1, higher is better $0.000 $0.024 $0.049 $0.073 $0.098 judge spend per 1,000 cases (USD) Production-weighted · OpenJev (Q2) — rules → small model, no judge: judge spend $0.00000 per 1,000 cases, block-only F1 0.21285 — matched or beaten on both, with at least one strict loss Production-weighted · DiffusionGemma (Q3) — rules → small model, no judge: judge spend $0.00000 per 1,000 cases, block-only F1 0.10717 — matched or beaten on both, with at least one strict loss Production-weighted · Jev 1.13.0 (Q2) — rules → small model, no judge: judge spend $0.00000 per 1,000 cases, block-only F1 0.23301 — undominated Production-weighted · no small model — rules → judge, no small model: judge spend $0.0923 per 1,000 cases, block-only F1 0.10980 — matched or beaten on both, with at least one strict loss Production-weighted · OpenJev (Q2) — rules → small → judge, one-sided: judge spend $0.0884 per 1,000 cases, block-only F1 0.06134 — matched or beaten on both, with at least one strict loss Production-weighted · OpenJev (Q2) — rules → small → judge, two-sided @0.05: judge spend $0.0417 per 1,000 cases, block-only F1 0.06209 — matched or beaten on both, with at least one strict loss Production-weighted · OpenJev (Q2) — rules → small → judge, two-sided @0.10: judge spend $0.0256 per 1,000 cases, block-only F1 0.06385 — matched or beaten on both, with at least one strict loss Production-weighted · OpenJev (Q2) — rules → small → judge, two-sided @0.20: judge spend $0.0147 per 1,000 cases, block-only F1 0.07235 — matched or beaten on both, with at least one strict loss Production-weighted · OpenJev (Q2) — rules → small → judge, two-sided @0.30: judge spend $0.00980 per 1,000 cases, block-only F1 0.08215 — matched or beaten on both, with at least one strict loss Production-weighted · DiffusionGemma (Q3) — rules → small → judge, one-sided: judge spend $0.0790 per 1,000 cases, block-only F1 0.06445 — matched or beaten on both, with at least one strict loss Production-weighted · DiffusionGemma (Q3) — rules → small → judge, two-sided @0.05: judge spend $0.0155 per 1,000 cases, block-only F1 0.06810 — matched or beaten on both, with at least one strict loss Production-weighted · DiffusionGemma (Q3) — rules → small → judge, two-sided @0.10: judge spend $0.0120 per 1,000 cases, block-only F1 0.06713 — matched or beaten on both, with at least one strict loss Production-weighted · DiffusionGemma (Q3) — rules → small → judge, two-sided @0.20: judge spend $0.00856 per 1,000 cases, block-only F1 0.06783 — matched or beaten on both, with at least one strict loss Production-weighted · DiffusionGemma (Q3) — rules → small → judge, two-sided @0.30: judge spend $0.00650 per 1,000 cases, block-only F1 0.06470 — matched or beaten on both, with at least one strict loss Production-weighted · Jev 1.13.0 (Q2) — rules → small → judge, one-sided: judge spend $0.0838 per 1,000 cases, block-only F1 0.07708 — matched or beaten on both, with at least one strict loss Production-weighted · Jev 1.13.0 (Q2) — rules → small → judge, two-sided @0.05: judge spend $0.0557 per 1,000 cases, block-only F1 0.07736 — matched or beaten on both, with at least one strict loss Production-weighted · Jev 1.13.0 (Q2) — rules → small → judge, two-sided @0.10: judge spend $0.0429 per 1,000 cases, block-only F1 0.07835 — matched or beaten on both, with at least one strict loss Production-weighted · Jev 1.13.0 (Q2) — rules → small → judge, two-sided @0.20: judge spend $0.0304 per 1,000 cases, block-only F1 0.08227 — matched or beaten on both, with at least one strict loss Production-weighted · Jev 1.13.0 (Q2) — rules → small → judge, two-sided @0.30: judge spend $0.0226 per 1,000 cases, block-only F1 0.08682 — matched or beaten on both, with at least one strict loss Jev 1.13.0 Q2 · rules → small
Data table
CorpusSmall modelPolicyjudge spend per 1,000 casesblock-only F1block FPRstanding
Broad comparisonOpenJev (Q2)rules → small model — no judge$0.00000 (no judge)0.702310.00385undominated
Broad comparisonDiffusionGemma (Q3)rules → small model — no judge$0.00000 (no judge)0.387650.01006matched or beaten on both, with at least one strict loss
Broad comparisonJev 1.13.0 (Q2)rules → small model — no judge$0.00000 (no judge)0.541530.00089matched or beaten on both, with at least one strict loss
Broad comparisonno small modelrules → judge — no small model$0.1940.712480.00680matched or beaten on both, with at least one strict loss
Broad comparisonOpenJev (Q2)rules → small → judge — one-sided$0.1750.731570.00592matched or beaten on both, with at least one strict loss
Broad comparisonOpenJev (Q2)rules → small → judge — two-sided @0.05$0.09290.731570.00592matched or beaten on both, with at least one strict loss
Broad comparisonOpenJev (Q2)rules → small → judge — two-sided @0.10$0.07620.732590.00562matched or beaten on both, with at least one strict loss
Broad comparisonOpenJev (Q2)rules → small → judge — two-sided @0.20$0.04920.735660.00473matched or beaten on both, with at least one strict loss
Broad comparisonOpenJev (Q2)rules → small → judge — two-sided @0.30$0.03090.737730.00414undominated
Broad comparisonDiffusionGemma (Q3)rules → small → judge — one-sided$0.1770.488960.01272matched or beaten on both, with at least one strict loss
Broad comparisonDiffusionGemma (Q3)rules → small → judge — two-sided @0.05$0.02290.488820.01094matched or beaten on both, with at least one strict loss
Broad comparisonDiffusionGemma (Q3)rules → small → judge — two-sided @0.10$0.01470.484750.01065matched or beaten on both, with at least one strict loss
Broad comparisonDiffusionGemma (Q3)rules → small → judge — two-sided @0.20$0.009630.467530.01065matched or beaten on both, with at least one strict loss
Broad comparisonDiffusionGemma (Q3)rules → small → judge — two-sided @0.30$0.006790.455740.01035matched or beaten on both, with at least one strict loss
Broad comparisonJev 1.13.0 (Q2)rules → small → judge — one-sided$0.1750.659700.00385matched or beaten on both, with at least one strict loss
Broad comparisonJev 1.13.0 (Q2)rules → small → judge — two-sided @0.05$0.1170.659700.00385matched or beaten on both, with at least one strict loss
Broad comparisonJev 1.13.0 (Q2)rules → small → judge — two-sided @0.10$0.09920.661680.00325matched or beaten on both, with at least one strict loss
Broad comparisonJev 1.13.0 (Q2)rules → small → judge — two-sided @0.20$0.06600.663660.00266matched or beaten on both, with at least one strict loss
Broad comparisonJev 1.13.0 (Q2)rules → small → judge — two-sided @0.30$0.04860.659640.00266matched or beaten on both, with at least one strict loss
Production-weightedOpenJev (Q2)rules → small model — no judge$0.00000 (no judge)0.212850.00924matched or beaten on both, with at least one strict loss
Production-weightedDiffusionGemma (Q3)rules → small model — no judge$0.00000 (no judge)0.107170.05351matched or beaten on both, with at least one strict loss
Production-weightedJev 1.13.0 (Q2)rules → small model — no judge$0.00000 (no judge)0.233010.00590undominated
Production-weightedno small modelrules → judge — no small model$0.09230.109800.09025matched or beaten on both, with at least one strict loss
Production-weightedOpenJev (Q2)rules → small → judge — one-sided$0.08840.061340.07038matched or beaten on both, with at least one strict loss
Production-weightedOpenJev (Q2)rules → small → judge — two-sided @0.05$0.04170.062090.06939matched or beaten on both, with at least one strict loss
Production-weightedOpenJev (Q2)rules → small → judge — two-sided @0.10$0.02560.063850.06465matched or beaten on both, with at least one strict loss
Production-weightedOpenJev (Q2)rules → small → judge — two-sided @0.20$0.01470.072350.05570matched or beaten on both, with at least one strict loss
Production-weightedOpenJev (Q2)rules → small → judge — two-sided @0.30$0.009800.082150.04671matched or beaten on both, with at least one strict loss
Production-weightedDiffusionGemma (Q3)rules → small → judge — one-sided$0.07900.064450.07260matched or beaten on both, with at least one strict loss
Production-weightedDiffusionGemma (Q3)rules → small → judge — two-sided @0.05$0.01550.068100.06691matched or beaten on both, with at least one strict loss
Production-weightedDiffusionGemma (Q3)rules → small → judge — two-sided @0.10$0.01200.067130.06568matched or beaten on both, with at least one strict loss
Production-weightedDiffusionGemma (Q3)rules → small → judge — two-sided @0.20$0.008560.067830.06370matched or beaten on both, with at least one strict loss
Production-weightedDiffusionGemma (Q3)rules → small → judge — two-sided @0.30$0.006500.064700.06242matched or beaten on both, with at least one strict loss
Production-weightedJev 1.13.0 (Q2)rules → small → judge — one-sided$0.08380.077080.04745matched or beaten on both, with at least one strict loss
Production-weightedJev 1.13.0 (Q2)rules → small → judge — two-sided @0.05$0.05570.077360.04725matched or beaten on both, with at least one strict loss
Production-weightedJev 1.13.0 (Q2)rules → small → judge — two-sided @0.10$0.04290.078350.04651matched or beaten on both, with at least one strict loss
Production-weightedJev 1.13.0 (Q2)rules → small → judge — two-sided @0.20$0.03040.082270.04374matched or beaten on both, with at least one strict loss
Production-weightedJev 1.13.0 (Q2)rules → small → judge — two-sided @0.30$0.02260.086820.04086matched or beaten on both, with at least one strict loss
The line joins the policies no other policy in the same panel matches or beats on both axes with at least one strict improvement. A hollow dot is matched or beaten on both, with at least one strict loss. On the broad comparison corpus 2 of 19 policies are undominated: OpenJev (Q2) two-sided @0.30 at 0.73773 block F1, OpenJev (Q2) rules → small at 0.70231 block F1. On the production-weighted corpus 1 of 19 policies are undominated: Jev 1.13.0 (Q2) rules → small at 0.23301 block F1. Policies with no judge sit on the left edge at $0. Source: outputs/deterministic-real/realdet-s2-openjev.json, outputs/deterministic-real/realdet-s3-openjev.json :: candidates[0].<policy>.{binary_block_only.f1, llm_invocation_rate}; judge price from outputs/s2/gemma4-c7.jsonl.meta.json :: prompt_tokens / cases / provider_calls (the Q0 judge run the cascade was scored from, which records no price) at the rate in outputs/s2/gemma4-q2.jsonl.meta.json :: estimated_usd / prompt_tokens (the Q2 judge run, the only one priced) and outputs/s3/gemma4-c7.jsonl.meta.json :: prompt_tokens / cases / provider_calls (the Q0 judge run the cascade was scored from, which records no price) at the rate in outputs/s2/gemma4-q2.jsonl.meta.json :: estimated_usd / prompt_tokens (the Q2 judge run, the only one priced); outputs/deterministic-real/realdet-s2-jev.json :: candidates[0] (jev-1.13.0/C7/I3/Q2); outputs/deterministic-real/realdet-s3-jev.json :: candidates[0] (jev-1.13.0/C7/I3/Q2).

The cascade as a flow

enteringdeterministic rulessmall modelLLM judgeends allowends confirmends block
Every case flowing through the three tiers to a final dispositionOf 3,817 cases the deterministic rules terminate 13, the small model decides 3,195 and 609 reach the judge. 2,900 end allow, 640 confirm and 277 block. Broad comparison, the shipped policy: rules, then OpenJev, then Gemma 4 on the uncertain band only, trusted-allow threshold 0.30. All 3,817 scorable cases enter at the left. Ribbon width is case count, and each arrow carries the measured rate of the traffic taking it. The merge is coloured by deciding tier; the final split is the pooled measured mix. See the note. Decisions in: 3,817 of 3,817 cases = 100.00% Decisions in 3,817 scorable cases Deterministic rules: 3,817 of 3,817 cases = 100.00% Deterministic rules terminate or pass Small model (OpenJev): 3,804 of 3,817 cases = 99.66% Small model (OpenJev) decide or escalate LLM judge (Gemma 4): 609 of 3,817 cases = 15.95% LLM judge (Gemma 4) 15.95% of cases reach it Final disposition what the agent actually gets ends allow: 2,900 of 3,817 cases = 75.98% allow 75.98% ends confirm: 640 of 3,817 cases = 16.77% confirm 16.77% ends block: 277 of 3,817 cases = 7.26% block 7.26% 13 terminate (0.34%) 3,804 pass (99.66%) 609 escalate (15.95%) 3,195 decided (83.70%)
Data table
Junctioncasesshare of all cases
all cases entering3,8171
terminated by the deterministic rules130.003406
passed through the rules to the small model3,8040.996594
decided by the small model, judge never called3,1950.837045
escalated to the judge6090.159549
ends allow2,9000.759759
ends confirm6400.167671
ends block2770.072570
The scorecard records which tier decided each case, and the final disposition mix, but not the two cross-tabulated — so the last band is drawn from the pooled measured mix. The rules terminate only 13 of 3,817 cases, and 8 of those are advisory confirms that capped a block the cascade would otherwise have issued. The judge sees 15.95% of traffic, which is the whole point of the two-sided band: one-sided routing sends 90.31% of cases to it for 0.00616 less block F1. Source: outputs/s2/policy-reanalysis-realdet.json :: compositions/realdet_short_circuit/cascade_tiers/two_tier_openjev_then_gemma.{decided_by, gemma_invocation_rate, counts}; final mix from outputs/deterministic-real/realdet-s2-openjev.json :: candidates[0].deterministic_then_system_one_then_llm_two_sided_0.30.three_way.confusion.

Review burden against attacks allowed through

OpenJevDiffusionGemmaJev 1.13.0rules → judge (reference)
Review burden against attacks allowed throughRaising the trusted-allow threshold walks each model down and to the left: less review work, more attacks through. The four measured thresholds are the four points on each path. Horizontal: share of cases ending confirm, which is what lands on a human or a second model. Vertical: share of the unsafe cases whose final disposition is allow. One path per model, four measured thresholds each, joined 0.05 to 0.30. Bottom-left is better on both. Broad comparison, 3 paths. 0% 10% 20% 30% 40% 50% 0% 5% 10% 16% 21% 26% confirm rate — share of cases left for review unsafe cases allowed through on the vertical axis OpenJev at allow≤0.05: confirm rate 39.088%, unsafe allowed 8.716%, block F1 0.73157 OpenJev at allow≤0.10: confirm rate 33.953%, unsafe allowed 9.404%, block F1 0.73259 OpenJev at allow≤0.20: confirm rate 23.946%, unsafe allowed 11.697%, block F1 0.73566 OpenJev at allow≤0.30: confirm rate 16.767%, unsafe allowed 13.991%, block F1 0.73773 OpenJev @0.30 DiffusionGemma at allow≤0.05: confirm rate 14.802%, unsafe allowed 16.514%, block F1 0.48882 DiffusionGemma at allow≤0.10: confirm rate 11.239%, unsafe allowed 18.349%, block F1 0.48475 DiffusionGemma at allow≤0.20: confirm rate 8.829%, unsafe allowed 20.642%, block F1 0.46753 DiffusionGemma at allow≤0.30: confirm rate 7.545%, unsafe allowed 22.248%, block F1 0.45574 DiffusionGemma @0.30 Jev 1.13.0 at allow≤0.05: confirm rate 43.987%, unsafe allowed 9.404%, block F1 0.65970 Jev 1.13.0 at allow≤0.10: confirm rate 42.337%, unsafe allowed 11.468%, block F1 0.66168 Jev 1.13.0 at allow≤0.20: confirm rate 33.036%, unsafe allowed 12.615%, block F1 0.66366 Jev 1.13.0 at allow≤0.30: confirm rate 26.906%, unsafe allowed 14.679%, block F1 0.65964 Jev 1.13.0 @0.30 rules → judge, no small model: confirm rate 43.070%, unsafe allowed 9.404%, block F1 0.71248 rules → judge, no small model
Data table
Modelallow thresholdconfirm rateunsafe allowed throughblock-only F1
OpenJev0.050.3908830.0871560.73157
OpenJev0.100.3395340.0940370.73259
OpenJev0.200.2394550.1169720.73566
OpenJev0.300.1676710.1399080.73773
DiffusionGemma0.050.1480220.1651380.48882
DiffusionGemma0.100.1123920.1834860.48475
DiffusionGemma0.200.0882890.2064220.46753
DiffusionGemma0.300.0754520.2224770.45574
Jev 1.13.00.050.4398740.0940370.65970
Jev 1.13.00.100.4233690.1146790.66168
Jev 1.13.00.200.3303640.1261470.66366
Jev 1.13.00.300.2690590.1467890.65964
rules → judge (reference)—0.4307050.0940370.71248
As the trusted-allow threshold rises each path moves left and up: less review, more leakage. Between 0.05 and 0.30 OpenJev gives up 22.32 points of review burden and takes on 5.28 points of extra leakage. DiffusionGemma leaks more than OpenJev at every threshold both were measured at, and their review-rate ranges do not overlap (7.55–14.80% against 16.77–39.09%), so there is no measured review level to compare them at. The closest measured pair is DiffusionGemma at 14.80% confirm leaking 16.51%, against OpenJev at 16.77% confirm leaking 13.99%. Source: outputs/deterministic-real/realdet-s2-openjev.json, outputs/deterministic-real/realdet-s2-diffgemma.json :: candidates[0].deterministic_then_system_one_then_llm_two_sided_* :: {review_rate, binary.confusion}; outputs/deterministic-real/realdet-s2-jev.json :: candidates[0] (jev-1.13.0/C7/I3/Q2).

Price it against your own traffic

The projection prices the judge only, from the measured rates: not your own tools, retries, in-line latency or the human cost of a confirm.

Your traffic0.30Measured at allow thresholds 0.05, 0.10, 0.20, 0.30 and at 88.58% and 99.10% benign; anything between is a straight line between measurements, anything outside is clamped. A case averaged 7.9408 decisions on the Broad comparison and 4.0857 on the Production-weighted corpus, if you count decisions rather than cases. With scripting off the tiles show 14,000 cases per day at 99.10% benign and threshold 0.30.
judge spend per month
$4.12
USD
confirms per day
1,357
cases sent for review
unsafe calls allowed through per day
5.7
final disposition allow
The interpolated rates the projection above is built from, and the constants. The judge price is a per-corpus constant and is interpolated on the same axis as the rates: $0.000194 per case on the Broad comparison, $0.0000923 per case on the Production-weighted. Each is that corpus’s own judge run priced at the $0.042 per million input tokens that outputs/s2/gemma4-q2.jsonl.meta.json records; the scored runs (outputs/s2/gemma4-c7.jsonl.meta.json, outputs/s3/gemma4-c7.jsonl.meta.json) carry token counts but no price of their own. A month is 30 days.
Interpolated ratevalue
share of cases that reach the judge0.106227
share of cases ending confirm0.096952
share of unsafe cases ending allow0.045249
cases per day, as entered14,000
benign share actually used (clamped to the measured range)0.9910

Serving cost and latency

Recorded provider spend over every run on this site is $23.95: $23.18 for hosted Jev 1.13.0 (551,834,284 input tokens at $0.042 per million, the one rate every priced manifest records) and $0.777 for the Gemma 4 judge, whose large-stage runs logged no cost, so its true Bedrock spend is higher by an unknown amount. The self-hosted models ran on one g6e.12xlarge (published on-demand rate $12.90 per hour); its uptime is not recorded, so no GPU total is stated.

The small model emits one token per call, so serving time is all prefill: there is no decode phase to speed up and input tokens are the whole cost. The latencies below come from a saturated GPU running a batch to completion; the long tail is queueing, and any in-line deployment needs a timeout policy.

Wall-clock latency under batch load

Wall-clock times per case. The self-hosted rows come from a saturated GPU running a batch to completion, so a single in-line call would see less; a hosted-API row is a network round trip and is not comparable with them.

p50 (median)p95p50 → p99 range
End-to-end wall-clock latency per caseEvery stage and model with a recorded latency block is drawn. Medians run from 0.40 to 19.44 seconds per case and the p99 tail to 252.7. The self-hosted rows reflect queueing on a saturated GPU; a hosted row is a network round trip. 0s 50s 100s 150s 200s 250s Median to p99 per case (bar), with p50 and p95 marked Broad comparison · OpenJev Broad comparison · OpenJev: p50 19.4s, p95 120.5s, p99 252.7s, max 434.1s, n=3,817 p99 253s Broad comparison · DiffusionGemma Broad comparison · DiffusionGemma: p50 3.7s, p95 23.9s, p99 48.0s, max 93.6s, n=3,817 p99 48s Broad comparison · Jev 1.13.0 Broad comparison · Jev 1.13.0: p50 1.4s, p95 7.4s, p99 14.5s, max 21.2s, n=3,817 p99 14s Production-weighted · OpenJev Production-weighted · OpenJev: p50 5.2s, p95 42.2s, p99 47.8s, max 365.4s, n=24,476 p99 48s Production-weighted · DiffusionGemma Production-weighted · DiffusionGemma: p50 1.4s, p95 10.6s, p99 13.6s, max 75.4s, n=24,476 p99 14s Production-weighted · Jev 1.13.0 Production-weighted · Jev 1.13.0: p50 0.4s, p95 3.1s, p99 3.4s, max 22.0s, n=24,476 p99 3s wall-clock seconds per case
Data table
Rundecisionsp50 (s)p95 (s)p99 (s)max (s)spendinput tokenserrors
Broad comparison · OpenJev3,81719.44120.52252.70434.1$1.9746,902,0850
Broad comparison · DiffusionGemma3,8173.6623.8747.9993.6$1.2930,791,7410
Broad comparison · Jev 1.13.03,8171.427.3514.4521.2$1.2128,859,8890
Production-weighted · OpenJev24,4765.1942.2347.75365.4$5.34127,247,45432
Production-weighted · DiffusionGemma24,4761.3610.6313.6075.4$4.0696,676,7290
Production-weighted · Jev 1.13.024,4760.403.073.4222.0$3.7288,682,7520
Source: outputs/deterministic-real/realdet-s2-openjev.json; outputs/deterministic-real/realdet-s2-diffgemma.json; outputs/deterministic-real/realdet-s2-jev.json; outputs/deterministic-real/realdet-s3-openjev.json; outputs/deterministic-real/realdet-s3-diffgemma.json; outputs/deterministic-real/realdet-s3-jev.json :: candidates[0].system_one.latency_ms / estimated_usd / input_tokens / errors.
The judge on the flex service tier: 200 cases, identical prompts. Source: outputs/flex/.
Metricdefaultflexchange
median per-case latency844.6 ms813.8 ms−3.6%
p95 per-case latency23,502.7 ms15,908.9 ms−32.3%
worst case122,211 ms33,560 ms−72.5%
tokens, errors993,812 / 3,038, 0993,812 / 3,038, 0identical
case decisions199 of 200 identical

Prefix caching does not pay for this workload. 82.14% of 1,008 measured prompts are shorter than one 784-token cache block. Padding prompts to the block boundary cost +86.68% input tokens. A finer cache match unit raised hit rates but slowed the server: on the 120-prompt probe by 34.7% serially and 15.6% at four concurrent requests, and on the 200-case benchmark workload by 0.72% of wall clock. Putting the static part of the prompt first costs nothing.

Recommendations

Twelve recommendations. Confidence: high means reproduced and gated; medium means measured once or with a stated confound.
#RecommendationEvidenceConfidence
1Enforce per tool call, not per session.2.7×–5.1× less benign work destroyed at identical recall (above).high
2Use a single-disposition question (Q2, or Q4) for anything that blocks; never the boolean-probe formats Q1 and Q3.Block-only F1 0.70231 against 0.03160 and 0.03604; Q1 and Q3 fail open on a malformed answer (Risks).high
3Set the escalation threshold per surface to save judge calls.1.0959 points fewer judge calls on one-shot actions at unchanged F1.qualified
4Treat a rule-engine confirm as a floor and keep evaluating.0.73773 → 0.75173 block-only F1, same false blocks and judge-call rate.high
5At production traffic, run rules → small model with no judge, once recommendation 6 is in place.0.08215 → 0.21285 block-only F1 at zero judge cost (above).high
6Make the disposition on a small-model failure configurable, defaulting to confirm, with a timeout.24 of 28 injected failure classes end with the call allowed (Risks).high
7Decide explicitly whether to accept the short-circuit loss of two-sided routing.Of the 1,346 cases the judge wanted to escalate and the small model allowed, 32 are graded unsafe (2.38%); two-sided routing discards at most that set (Risks).medium
8Author the mined deterministic rules, starting with shell-history tampering.96.2% of the small model’s confident blocks match no rule; that cluster is grade A throughout (20 events).medium
9Fix the two provenance defects in the dataset lock.One disabled source was normalised into the intent corpus; one entry pins a paper, not a revision (Method).high
10Do not use the serves-the-request score (Lane B) as a standalone gate.It flags 0.17887 of benign tool calls and 39/40 benign coding sessions (Risks).high
11Do not invest in prefix caching or prompt padding.+86.68% input tokens for padding; slower serving with a finer match unit (above).qualified
12Do not reinstate Von 1.0.1.Block-only F1 0.76786 against 0.71545 for blocking everything on the same pilot, at a 38.8s–44.4s median per case on CPU.high

What would change these numbers

AssumptionWhat breaks it
The judge price, $0.194 per 1,000 cases judgedA different context length, service tier or provider price.
The benign share, between 88.58% and 99.10%A third mix; nothing between or outside the two was run.
A confirm stops nothingIf your confirm path is enforced by a reliable second decider, any-intervention F1F1 counting a confirm as a catch as well as a block. It measures triage; it overstates what a guard actually stops. becomes the relevant score and the order of policies changes.
Recall on proven compromisesEven the best model blocks only 0.5778 of proof-verified compromises (Intent). No setting here makes the guard sufficient on its own.

evaluation-only never-train
Nothing derived here is approved for training, synthetic generation, teacher context, distillation or redistribution. The 13 source datasets are public and linked on reproduce; the analysis artifacts behind each number are held privately and are available on request.