Reproduce

Every figure comes from a JSON or text artifact written by a named script, over a prediction file with a recorded hash, at a fixed seed. The analysis artifacts are held privately and are available on request; the source datasets are public.

Sources

The 13 sources that supply rows are public and reachable without a token. Each is download-only: the table links them and this site republishes none of their rows. Licences in use: Apache-2.0, CC-BY-4.0, CC-BY-NC-4.0, MIT. coding-agent security benchmark is CC-BY-NC-4.0 and its 279 rows must not be used for any commercial purpose.

The 13 attributed sources that supply rows, summing to 413,180. Every one is public and reachable without a token. Licence, URL and pinned revision come from benchmarks/datasets.lock.json; row counts from outputs/source-catalog-v2.json. Each is download-only: this site links them and never republishes their rows.
Sourcerowslicencegradeswhat it holds
Nemotron agentic terminal pivot
nvidia/Nemotron-RL-Agentic-Terminal-Pivot-v1
195,768CC-BY-4.0D 195,768benign verifier-passing terminal trajectories, the bulk of the benign mass
ResearchArena trajectories
aisa-group/ResearchArena-Trajectories
66,469Apache-2.0D 15,077 / E 51,392coding-agent trajectories
AgentHazard
Yunhao-Feng/AgentHazard
65,293MITA 13 / E 65,280harmful-scenario action and stateful cases
MCP-hunt agent traces
lihaonan0716/mcphunt-agent-traces
40,709CC-BY-4.0A 1 / B 768 / D 26,163 / E 13,777executed MCP attack chains plus paired benign
agentic red-team (synthetic)
hf-agentic-red-team
18,265Apache-2.0C 17 / E 18,248synthetic command and chain positives; 30 grade-A rows here were revoked
cochise
andreashappe/cochise
9,660MITA 290 / D 4 / E 9,366290 of the 333 remaining grade-A rows
DTAP-Bench agent trajectories
AI-Secure/DTap-Bench-Agent-Trajectories
6,540Apache-2.0D 4,692 / E 1,848benign and hard-negative tool trajectories with real arguments
ctrl-dataset / monitoringbench
neur26anonsub/ctrldataset2026
4,883CC-BY-4.0B 2,441 / C 2,442environment-verified attacks with real arguments
quadrat IPI model eval
mihail-gribov/quadrat-ipi-model-eval
2,120Apache-2.0D 850 / E 1,270prompt-injection evaluation
enigma-agent trajectories
enigma-agent/trajectories
1,418MITE 1,418all out of scope
sentinel-flow
sentinel-flow
900Apache-2.0B 486 / C 4 / D 410synthetic conformance with source–sink lineage
agentic red-team benchmark
agentic-redteam-benchmark
876CC-BY-4.0C 438 / D 438red-team scenarios
coding-agent security benchmark
rogue-coding-agent-security
279CC-BY-NC-4.0 non-commercialA 29 / C 75 / D 65 / E 110split across two normalised corpora

9 of the sources are Hugging Face datasets and can be previewed below in the Hugging Face viewer, with the pinned revision beside each; the other 4 are on GitHub. Every preview is collapsed and loads nothing from huggingface.co until opened, and the rows it shows are served by Hugging Face from the source repository, not republished here.

Nemotron agentic terminal pivot195,768 rowsCC-BY-4.0

nvidia/Nemotron-RL-Agentic-Terminal-Pivot-v1 · benign verifier-passing terminal trajectories, the bulk of the benign mass
open the source · pinned at eaef26944643

Preview the rows in the Hugging Face dataset viewer
ResearchArena trajectories66,469 rowsApache-2.0

aisa-group/ResearchArena-Trajectories · coding-agent trajectories
open the source · pinned at 241e456ba88b

Preview the rows in the Hugging Face dataset viewer
AgentHazard65,293 rowsMIT

Yunhao-Feng/AgentHazard · harmful-scenario action and stateful cases
open the source · pinned at 786147ad768f

Preview the rows in the Hugging Face dataset viewer
MCP-hunt agent traces40,709 rowsCC-BY-4.0

lihaonan0716/mcphunt-agent-traces · executed MCP attack chains plus paired benign
open the source · pinned at c4c69322bf9e

Preview the rows in the Hugging Face dataset viewer
agentic red-team (synthetic)18,265 rowsApache-2.0

hf-agentic-red-team · synthetic command and chain positives; 30 grade-A rows here were revoked
open the source · pinned at 3ed8b32aaed0

Preview the rows in the Hugging Face dataset viewer
cochise9,660 rowsMIT

andreashappe/cochise · 290 of the 333 remaining grade-A rows
open the source · pinned at 3abdb11f577d

GitHub-hosted: link only.

DTAP-Bench agent trajectories6,540 rowsApache-2.0

AI-Secure/DTap-Bench-Agent-Trajectories · benign and hard-negative tool trajectories with real arguments
open the source · pinned at 836caf2fdd78

Preview the rows in the Hugging Face dataset viewer
ctrl-dataset / monitoringbench4,883 rowsCC-BY-4.0

neur26anonsub/ctrldataset2026 · environment-verified attacks with real arguments
open the source · pinned at 6c72cf198833

Preview the rows in the Hugging Face dataset viewer
quadrat IPI model eval2,120 rowsApache-2.0

mihail-gribov/quadrat-ipi-model-eval · prompt-injection evaluation
open the source · pinned at 03e496ba1af2

GitHub-hosted: link only.

enigma-agent trajectories1,418 rowsMIT

enigma-agent/trajectories · all out of scope
open the source · pinned at 431dda6896b2

GitHub-hosted: link only.

sentinel-flow900 rowsApache-2.0

sentinel-flow · synthetic conformance with source–sink lineage
open the source · pinned at f3b8d86ab1b2

Preview the rows in the Hugging Face dataset viewer
agentic red-team benchmark876 rowsCC-BY-4.0

agentic-redteam-benchmark · red-team scenarios
open the source · pinned at 12ddc82333e5

GitHub-hosted: link only.

coding-agent security benchmark279 rowsCC-BY-NC-4.0

rogue-coding-agent-security · split across two normalised corpora
open the source · pinned at bf7ff748d80c

Preview the rows in the Hugging Face dataset viewer

What is withheld, and the one carve-out

The proof-backed corpus is aggregate-only. It carries redistribution: aggregate-only in the dataset lock and license: unresolved in its own manifest (the upstream repository ships no LICENSE file), and publication is restricted to local evaluation. This site reports its rates, counts and intervals and publishes no case, request, tool call or description from it.

The carve-out. One table publishes per-case structure: the disagreement queue on Risks. Each of its 2,133 rows is 9 small integers indexing code tables printed beside it (truth grade, surface, event count, source dataset, four votes and a verdict), and the rows collapse to 500 distinct tuples. No case_id, request, tool call, argument value or free-form rationale appears in it, and no row comes from the aggregate-only or local-evaluation-only corpora. Coded contingency data carrying no identifier and no text is permitted under the rule; this paragraph states that permission so the rule and the payload agree.

Instantiated prompts are withheld: each embeds a corpus row. Method publishes the templates and one synthetic instantiation.

Code paths

Everything below is in feat/system-one-benchmarks, and pinned at d2ae73f32736, which is the tree these figures were produced from. The default branch does not carry this work: of the 11 paths below, only benchmarks/datasets.lock.json resolves on main, and there it is a different file. Each row links the branch copy and the pinned copy; a branch moves.

PathWhat it does
benchmarks/system_one/contexts-v1.json
at d2ae73f32736
10 context recipes (C0…CD) and production byte bounds.
benchmarks/system_one/questions-v1.json
at d2ae73f32736
The first question set, kept for the pilot arms that were run against it.
benchmarks/system_one/questions-v2.json
at d2ae73f32736
5 question formats (Q0…Q4), 4 instruction variants (I0…I3), verbatim prompt text.
benchmarks/schema/system-one-prediction-v1.schema.json
at d2ae73f32736
Prediction row schema with the enum of valid context variants.
benchmarks/scripts/benchmark_run_system_one.py
at d2ae73f32736
Runner. derive_action() holds the Q1/Q3 answer-type guard; validate_resume_prefix() the resume validator behind the tamper experiment.
benchmarks/scripts/benchmark_score_system_one.py
at d2ae73f32736
Cascade scorer. Produces the binary / binary_block_only / three_way blocks and all policy lenses.
benchmarks/scripts/score_intent_separation.py
at d2ae73f32736
Separation scorer. Unmodified across every intent analysis, including the AgentDojo control.
benchmarks/scripts/benchmark_inventory_system_one_sources.py
at d2ae73f32736
Truth-grade assignment (truth_grade()) and family-identity resolution (family_authority()).
benchmarks/scripts/benchmark_normalize_agenttrace.py
at d2ae73f32736
One of the per-corpus adapters; three were rewritten to carry the user’s request through normalisation.
benchmarks/datasets.lock.json
at d2ae73f32736
91 pinned sources with licence status and redistribution terms; 14 enabled: false. Linked at main, which is the copy this build reads: the branch copy is a different file with 11 disabled.
benchmarks/coverage-report.mapping.json
at d2ae73f32736
Per-source intended use and stated limitations.

Seeds and model revisions

Randomness

SettingValue
seed, everywhere741983
bootstrap replicates2,000
bootstrap unitfamily (strata.split_group)
paired comparisonscommon random numbers
single-proportion intervalsWilson 95%

Model revisions

ComponentRevision or identifier
OpenJev5ec9e5fd2f80a6fff386779b1e5ac7e389971889
DiffusionGemmadiffusiongemma-26B-A4B-it-FP8-dynamic
Jev 1.13.0jev-1.13.0
Gemma 4 judgegoogle.gemma-4-26b-a4b
Von 1.0.1f6b268ff47b449b688a8052dfb3c37c9518b18f1
label / adjudicator modelopenai.gpt-oss-120b-1:0
rule engine, at measurementa4e50bcdea829447bc8c97e9504de1e6bf379946

Von 1.0.1 would not answer under the library versions its SDK declared (80 of 80 requests failed with HTTP 422), so the environment was upgraded (transformers 4.57.6 → 5.17.0) and the artifact left as pinned.

Settled-file rules

  1. Gate on completion and hash. A prediction file is read only if its .meta.json says complete: true and its on-disk SHA-256 matches the recorded prediction_sha256, or if it is a recorded, byte-reproduced merge of files that each pass that rule.
  2. Reproduce before extending. Every analysis re-derives at least one published number through its own code path before emitting a new one.
  3. New paths only. No existing scorecard, prediction file, manifest or scorer was modified by any analysis.
  4. Publish the integrity counters: row counts, error rows, duplicate (case_id, event_index) pairs, rows missing probabilities and the case-set intersection between arms.

Worked example: the intent reversal

  1. Take the corpus and the eight prediction arms from outputs/intent-real/. Check cases.jsonl hashes to 85dfc730d1337bb3… and that all eight manifests report one cases_sha256.
  2. Per arm, verify complete: true and the on-disk SHA-256. Abort if either fails.
  3. Run benchmarks/scripts/score_intent_separation.py unmodified, with --lens both --lane both --per-event --bootstrap 2000 --seed 741983 --question Q2 --context {C0|C7}.
  4. Confirm the control: the same script, flags and seed on AgentDojo give -0.0208 (C0) and -0.1246 (C7) for Jev 1.13.0 on the lead cell (grade-A lane, block-only, per case, sep_vs_resisted). If the control does not reproduce the old negative, stop.
  5. Read the lead cell on the proof-backed corpus. Expect +0.2667 for OpenJev at C0 and +0.5750 at C7.
  6. The within-family contrast covers 325 families; all 32 cells should be positive with intervals excluding zero.

Site build and gates

The build writes _build-figures.json beside the pages: every substituted value, the artifacts read and the assertion count.

Experiment ledger

Every experiment behind this site, with its scale and standing. Negative results carry equal weight.
ExperimentScaleResultStatusWhere
Model and question-format screen (S1)200 scenariosCandidates culled; screening kills candidates, it does not rank survivors.settledMethod
Small model ahead of the judge (S2)4,277 scenariosBlock-only F1 0.71248 → 0.73773, judge calls 100% → 15.95%.settledDeploy
Same cascade at the production benign share (S3)24,476 scenariosCascade 0.08215 against 0.21285 with no judge.negativeDeploy
Every model at a shared false-block budget3,817 cases, 12 modelsopen-jev-qwen-27b first, OpenJev second.settledWhich model
Intent on AgentDojo651 matched cases, 8 runsNegative in all eight.supersededIntent
Intent on proven compromises3,402 cases, 405 proof-verified32 of 32 cells positive.settledIntent
Lane B as a gate40 benign sessionsFlags 39/40 sessions.negativeRisks
Benign coding traffic1,543 events0.01944 (OpenJev) to 0.33117 (Jev 1.13.0) flagged.settledRisks
Question format30,310 decisions × 3Only Q2 beats the judge inside the cascade.settledMethod
Judge re-run on the challenger’s question30,310 callsOn the challenger’s question the judge scores 0.64793 against 0.71248 on its own; the cascade’s advantage over it widens from +0.03925 to +0.10380.settled—
Per-call against per-session enforcement4 model/corpus pairs2.7×–5.1× less benign work per call.settledDeploy
Tier order4 orderingsStronger model first.settledDeploy
Per-surface threshold1,482 + 2,795 casesSaves judge calls, not detection.settledDeploy
Disagreement adjudication2,133 cases32 of 1,346 short-circuited cases graded unsafe.qualifiedRisks
Rule mining1,118 + 378 confident blocks96.2% match no rule; top candidate is grade E.qualifiedDeploy
Real rule engine against an all-allow stand-in4 stage/model pairsThe composition, not the tier, moves the cascade figures.restatedbelow
Replay determinism3 replays × 1,519 eventsOpenJev 0 flips; DiffusionGemma and Jev 1.13.0 unstable on flagged events.qualifiedRisks
Fault injection28 classes, 11 tampers24 of 28 end in allow; 5 of 11 tampers accepted.negativeRisks
Prefix-cache economics2,352 callsNo saving; padding costs +86.68% tokens.negativeDeploy
Flex service tier200 casesp95 −32.3%, identical tokens.settledDeploy
Von 1.0.1 re-measured12,152 requests0.76786 against a 0.71545 block-everything floor.settledWhich model
Dataset normalisation audit18 corporaThe request field was empty in every row.settledMethod

Restated figures

Where an earlier figure disagreed with the artifact, the artifact is the record.

Earlier claimArtifact
Intent deviation does not work; separation negative for every model. True on AgentDojo, false on proof-backed data: 32 of 32 cells positive.
Broad-comparison cascade block-only F1 0.75173. That figure was scored against an all-allow stand-in for the rule tier. With the real rule engine the cascade as it runs today scores 0.73773; escalating a rule confirm recovers exactly 0.75173, which is why that figure is the recommendation’s.
Production-weighted cascade block-only F1 0.12854. 0.08215 with the real rule engine; the stand-in figure inverts the finding.
Hosted Jev’s cascade figures excluded as scored on the stand-in tier. False: its tier file is byte-identical to the real rule-engine predictions (9d0df1e6b00f…). The exclusion was inferred from OpenJev’s file in the same directory.
Escalate-on-confirm restores 0.75173 in all orderings. In one ordering; the others reach 0.704385, 0.495208 and 0.465798.
DiffusionGemma leads on any-intervention F1. That was its Q3 cell ranked against other models’ Q2 cells. At the shared format it scores 0.26792 block-only and 0.18527 any-intervention.
A leaderboard ranked at each model’s shipped threshold. The ranking is now at one shared false-block budget; shipped F1 is a column.
The adjudicator allows 89.45% of the short-circuited class. 89.45% is the modal disagreement class (n = 1,337); the short-circuit-discard class (n = 1,346) is 89.08%.
341,828 normalised rows had no user request. That is an adapter docstring’s snapshot over 12 corpora; the catalogs record 413,180 over 18.
587 new grade-A families against a prior 16. 405 new proof-verified families (802 cases before de-duplication) against 333 grade-A rows inventory-wide, 16 in the development split.
All 21 fault classes produce allow. 28 classes were injected; 24 end in allow, 3 of them with no error.
~700,000 provider requests at a 0.0065% error rate. 1,808,873 requests across 168 manifests; 33 errors, 0.00182%.
A $0.0491 per million planning price for the hosted API. Every priced manifest records $0.042 per million input tokens.
16.67% is a structural prefix-cache ceiling; identical replay hits 89.06%. 16.67% is exact on a 20-case workload and the same arm scored lower on 200 cases; 89.06% is the 112-token cache unit only, 22.03% at the default 784.
Background-subtracted cache shares 25.89% → 44.65% → 62.69% for the prompt-layout arms. The subtraction is persisted in no artifact; the raw window shares are 11.97%, 18.24% and 33.00%.
~10% of flagged events are unstable across replays. Per model: 0.00% for OpenJev (0 of 143 flagged events); 17.24% for DiffusionGemma (10 of 58 flagged events); 13.82% for Jev 1.13.0 (21 of 152 flagged events).

evaluation-only never-train
Nothing derived here is approved for training, synthetic generation, teacher context, distillation or redistribution. The 13 source datasets are public and linked on reproduce; the analysis artifacts behind each number are held privately and are available on request.

Upstream defects reported: vllm-project/vllm#58015, wfzyx/von#11.