Reproduce
Every figure comes from a JSON or text artifact written by a named script, over a prediction file with a recorded hash, at a fixed seed. The analysis artifacts are held privately and are available on request; the source datasets are public.
Sources
The 13 sources that supply rows are public and reachable without a token. Each is
download-only: the table links them and this site republishes none of their rows.
Licences in use: Apache-2.0, CC-BY-4.0, CC-BY-NC-4.0, MIT. coding-agent security benchmark is CC-BY-NC-4.0 and its
279 rows must not be used for any commercial purpose.
| Source | rows | licence | grades | what it holds |
|---|---|---|---|---|
Nemotron agentic terminal pivotnvidia/Nemotron-RL-Agentic-Terminal-Pivot-v1 | 195,768 | CC-BY-4.0 | D 195,768 | benign verifier-passing terminal trajectories, the bulk of the benign mass |
ResearchArena trajectoriesaisa-group/ResearchArena-Trajectories | 66,469 | Apache-2.0 | D 15,077 / E 51,392 | coding-agent trajectories |
AgentHazardYunhao-Feng/AgentHazard | 65,293 | MIT | A 13 / E 65,280 | harmful-scenario action and stateful cases |
MCP-hunt agent traceslihaonan0716/mcphunt-agent-traces | 40,709 | CC-BY-4.0 | A 1 / B 768 / D 26,163 / E 13,777 | executed MCP attack chains plus paired benign |
agentic red-team (synthetic)hf-agentic-red-team | 18,265 | Apache-2.0 | C 17 / E 18,248 | synthetic command and chain positives; 30 grade-A rows here were revoked |
cochiseandreashappe/cochise | 9,660 | MIT | A 290 / D 4 / E 9,366 | 290 of the 333 remaining grade-A rows |
DTAP-Bench agent trajectoriesAI-Secure/DTap-Bench-Agent-Trajectories | 6,540 | Apache-2.0 | D 4,692 / E 1,848 | benign and hard-negative tool trajectories with real arguments |
ctrl-dataset / monitoringbenchneur26anonsub/ctrldataset2026 | 4,883 | CC-BY-4.0 | B 2,441 / C 2,442 | environment-verified attacks with real arguments |
quadrat IPI model evalmihail-gribov/quadrat-ipi-model-eval | 2,120 | Apache-2.0 | D 850 / E 1,270 | prompt-injection evaluation |
enigma-agent trajectoriesenigma-agent/trajectories | 1,418 | MIT | E 1,418 | all out of scope |
sentinel-flowsentinel-flow | 900 | Apache-2.0 | B 486 / C 4 / D 410 | synthetic conformance with source–sink lineage |
agentic red-team benchmarkagentic-redteam-benchmark | 876 | CC-BY-4.0 | C 438 / D 438 | red-team scenarios |
coding-agent security benchmarkrogue-coding-agent-security | 279 | CC-BY-NC-4.0 non-commercial | A 29 / C 75 / D 65 / E 110 | split across two normalised corpora |
9 of the sources are Hugging Face datasets and can be previewed below in the Hugging Face viewer, with the pinned revision beside each; the other 4 are on GitHub. Every preview is collapsed and loads nothing from huggingface.co until opened, and the rows it shows are served by Hugging Face from the source repository, not republished here.
nvidia/Nemotron-RL-Agentic-Terminal-Pivot-v1 · benign verifier-passing terminal trajectories, the bulk of the benign mass
open the source · pinned at eaef26944643
Preview the rows in the Hugging Face dataset viewer
aisa-group/ResearchArena-Trajectories · coding-agent trajectories
open the source · pinned at 241e456ba88b
Preview the rows in the Hugging Face dataset viewer
Yunhao-Feng/AgentHazard · harmful-scenario action and stateful cases
open the source · pinned at 786147ad768f
Preview the rows in the Hugging Face dataset viewer
lihaonan0716/mcphunt-agent-traces · executed MCP attack chains plus paired benign
open the source · pinned at c4c69322bf9e
Preview the rows in the Hugging Face dataset viewer
hf-agentic-red-team · synthetic command and chain positives; 30 grade-A rows here were revoked
open the source · pinned at 3ed8b32aaed0
Preview the rows in the Hugging Face dataset viewer
andreashappe/cochise · 290 of the 333 remaining grade-A rows
open the source · pinned at 3abdb11f577d
GitHub-hosted: link only.
AI-Secure/DTap-Bench-Agent-Trajectories · benign and hard-negative tool trajectories with real arguments
open the source · pinned at 836caf2fdd78
Preview the rows in the Hugging Face dataset viewer
neur26anonsub/ctrldataset2026 · environment-verified attacks with real arguments
open the source · pinned at 6c72cf198833
Preview the rows in the Hugging Face dataset viewer
mihail-gribov/quadrat-ipi-model-eval · prompt-injection evaluation
open the source · pinned at 03e496ba1af2
GitHub-hosted: link only.
enigma-agent/trajectories · all out of scope
open the source · pinned at 431dda6896b2
GitHub-hosted: link only.
sentinel-flow · synthetic conformance with source–sink lineage
open the source · pinned at f3b8d86ab1b2
Preview the rows in the Hugging Face dataset viewer
agentic-redteam-benchmark · red-team scenarios
open the source · pinned at 12ddc82333e5
GitHub-hosted: link only.
rogue-coding-agent-security · split across two normalised corpora
open the source · pinned at bf7ff748d80c
Preview the rows in the Hugging Face dataset viewer
What is withheld, and the one carve-out
The proof-backed corpus is aggregate-only. It carries
redistribution: aggregate-only in the dataset lock and license:
unresolved in its own manifest (the upstream repository ships no LICENSE file), and
publication is restricted to local evaluation. This site reports its rates, counts and intervals
and publishes no case, request, tool call or description from it.
The carve-out. One table publishes per-case structure: the disagreement
queue on Risks. Each of its 2,133 rows is
9 small integers indexing code tables printed beside it (truth grade, surface,
event count, source dataset, four votes and a verdict), and the rows collapse to
500 distinct tuples. No case_id, request, tool call, argument value
or free-form rationale appears in it, and no row comes from the aggregate-only or
local-evaluation-only corpora. Coded contingency data carrying no identifier and no text is
permitted under the rule; this paragraph states that permission so the rule and the payload
agree.
Instantiated prompts are withheld: each embeds a corpus row. Method publishes the templates and one synthetic instantiation.
Code paths
Everything below is in feat/system-one-benchmarks, and pinned at d2ae73f32736, which is the tree these figures were produced from. The default branch does not carry this work: of the 11 paths below, only benchmarks/datasets.lock.json resolves on main, and there it is a different file. Each row links the branch copy and the pinned copy; a branch moves.
| Path | What it does |
|---|---|
benchmarks/system_one/contexts-v1.jsonat d2ae73f32736 | 10 context recipes (C0…CD) and production byte bounds. |
benchmarks/system_one/questions-v1.jsonat d2ae73f32736 | The first question set, kept for the pilot arms that were run against it. |
benchmarks/system_one/questions-v2.jsonat d2ae73f32736 | 5 question formats (Q0…Q4), 4 instruction variants (I0…I3), verbatim prompt text. |
benchmarks/schema/system-one-prediction-v1.schema.jsonat d2ae73f32736 | Prediction row schema with the enum of valid context variants. |
benchmarks/scripts/benchmark_run_system_one.pyat d2ae73f32736 | Runner. derive_action() holds the Q1/Q3 answer-type guard; validate_resume_prefix() the resume validator behind the tamper experiment. |
benchmarks/scripts/benchmark_score_system_one.pyat d2ae73f32736 | Cascade scorer. Produces the binary / binary_block_only / three_way blocks and all policy lenses. |
benchmarks/scripts/score_intent_separation.pyat d2ae73f32736 | Separation scorer. Unmodified across every intent analysis, including the AgentDojo control. |
benchmarks/scripts/benchmark_inventory_system_one_sources.pyat d2ae73f32736 | Truth-grade assignment (truth_grade()) and family-identity resolution (family_authority()). |
benchmarks/scripts/benchmark_normalize_agenttrace.pyat d2ae73f32736 | One of the per-corpus adapters; three were rewritten to carry the user’s request through normalisation. |
benchmarks/datasets.lock.jsonat d2ae73f32736 | 91 pinned sources with licence status and redistribution terms; 14 enabled: false. Linked at main, which is the copy this build reads: the branch copy is a different file with 11 disabled. |
benchmarks/coverage-report.mapping.jsonat d2ae73f32736 | Per-source intended use and stated limitations. |
Seeds and model revisions
Randomness
| Setting | Value |
|---|---|
| seed, everywhere | 741983 |
| bootstrap replicates | 2,000 |
| bootstrap unit | family (strata.split_group) |
| paired comparisons | common random numbers |
| single-proportion intervals | Wilson 95% |
Model revisions
| Component | Revision or identifier |
|---|---|
| OpenJev | 5ec9e5fd2f80a6fff386779b1e5ac7e389971889 |
| DiffusionGemma | diffusiongemma-26B-A4B-it-FP8-dynamic |
| Jev 1.13.0 | jev-1.13.0 |
| Gemma 4 judge | google.gemma-4-26b-a4b |
| Von 1.0.1 | f6b268ff47b449b688a8052dfb3c37c9518b18f1 |
| label / adjudicator model | openai.gpt-oss-120b-1:0 |
| rule engine, at measurement | a4e50bcdea829447bc8c97e9504de1e6bf379946 |
Von 1.0.1 would not answer under the library versions its SDK declared (80 of 80 requests failed with HTTP 422), so the environment was upgraded (transformers 4.57.6 → 5.17.0) and the artifact left as pinned.
Settled-file rules
- Gate on completion and hash. A prediction file is read only if its
.meta.jsonsayscomplete: trueand its on-disk SHA-256 matches the recordedprediction_sha256, or if it is a recorded, byte-reproduced merge of files that each pass that rule. - Reproduce before extending. Every analysis re-derives at least one published number through its own code path before emitting a new one.
- New paths only. No existing scorecard, prediction file, manifest or scorer was modified by any analysis.
- Publish the integrity counters: row counts, error rows, duplicate
(case_id, event_index)pairs, rows missing probabilities and the case-set intersection between arms.
Worked example: the intent reversal
- Take the corpus and the eight prediction arms from
outputs/intent-real/. Checkcases.jsonlhashes to85dfc730d1337bb3…and that all eight manifests report onecases_sha256. - Per arm, verify
complete: trueand the on-disk SHA-256. Abort if either fails. - Run
benchmarks/scripts/score_intent_separation.pyunmodified, with--lens both --lane both --per-event --bootstrap 2000 --seed 741983 --question Q2 --context {C0|C7}. - Confirm the control: the same script, flags and seed on AgentDojo give
-0.0208(C0) and-0.1246(C7) for Jev 1.13.0 on the lead cell (grade-A lane, block-only, per case,sep_vs_resisted). If the control does not reproduce the old negative, stop. - Read the lead cell on the proof-backed corpus. Expect +0.2667 for OpenJev at C0 and +0.5750 at C7.
- The within-family contrast covers 325 families; all 32 cells should be positive with intervals excluding zero.
Site build and gates
- 1152 assertion checks, each a (file, key path, expected value) triple. A mismatch prints and writes nothing.
- Reading precision. Every figure is published rounded, never above six significant digits and money to cents; the artifact’s exact decimal is kept in the page (hover it, or use Exact values in the header). The verifier fails a page with any longer visible number or any exact value that does not round-trip.
- Structure and links. Every internal link must resolve to a page and an anchor; unresolved template tokens, overflowing chart labels and a tooltip inside a heading all abort the build.
- Payload guard. Before upload every file is scanned against the corpora themselves for case ids and verbatim text, plus CJK and credential patterns; a hit aborts.
- No external assets. Inline SVG, one inlined stylesheet, inline script the pages work without; the only embeds are the collapsed Hugging Face viewers above.
The build writes _build-figures.json beside the pages: every substituted value, the
artifacts read and the assertion count.
Experiment ledger
| Experiment | Scale | Result | Status | Where |
|---|---|---|---|---|
| Model and question-format screen (S1) | 200 scenarios | Candidates culled; screening kills candidates, it does not rank survivors. | settled | Method |
| Small model ahead of the judge (S2) | 4,277 scenarios | Block-only F1 0.71248 → 0.73773, judge calls 100% → 15.95%. | settled | Deploy |
| Same cascade at the production benign share (S3) | 24,476 scenarios | Cascade 0.08215 against 0.21285 with no judge. | negative | Deploy |
| Every model at a shared false-block budget | 3,817 cases, 12 models | open-jev-qwen-27b first, OpenJev second. | settled | Which model |
| Intent on AgentDojo | 651 matched cases, 8 runs | Negative in all eight. | superseded | Intent |
| Intent on proven compromises | 3,402 cases, 405 proof-verified | 32 of 32 cells positive. | settled | Intent |
| Lane B as a gate | 40 benign sessions | Flags 39/40 sessions. | negative | Risks |
| Benign coding traffic | 1,543 events | 0.01944 (OpenJev) to 0.33117 (Jev 1.13.0) flagged. | settled | Risks |
| Question format | 30,310 decisions × 3 | Only Q2 beats the judge inside the cascade. | settled | Method |
| Judge re-run on the challenger’s question | 30,310 calls | On the challenger’s question the judge scores 0.64793 against 0.71248 on its own; the cascade’s advantage over it widens from +0.03925 to +0.10380. | settled | — |
| Per-call against per-session enforcement | 4 model/corpus pairs | 2.7×–5.1× less benign work per call. | settled | Deploy |
| Tier order | 4 orderings | Stronger model first. | settled | Deploy |
| Per-surface threshold | 1,482 + 2,795 cases | Saves judge calls, not detection. | settled | Deploy |
| Disagreement adjudication | 2,133 cases | 32 of 1,346 short-circuited cases graded unsafe. | qualified | Risks |
| Rule mining | 1,118 + 378 confident blocks | 96.2% match no rule; top candidate is grade E. | qualified | Deploy |
| Real rule engine against an all-allow stand-in | 4 stage/model pairs | The composition, not the tier, moves the cascade figures. | restated | below |
| Replay determinism | 3 replays × 1,519 events | OpenJev 0 flips; DiffusionGemma and Jev 1.13.0 unstable on flagged events. | qualified | Risks |
| Fault injection | 28 classes, 11 tampers | 24 of 28 end in allow; 5 of 11 tampers accepted. | negative | Risks |
| Prefix-cache economics | 2,352 calls | No saving; padding costs +86.68% tokens. | negative | Deploy |
| Flex service tier | 200 cases | p95 −32.3%, identical tokens. | settled | Deploy |
| Von 1.0.1 re-measured | 12,152 requests | 0.76786 against a 0.71545 block-everything floor. | settled | Which model |
| Dataset normalisation audit | 18 corpora | The request field was empty in every row. | settled | Method |
Restated figures
Where an earlier figure disagreed with the artifact, the artifact is the record.
| Earlier claim | Artifact |
|---|---|
| Intent deviation does not work; separation negative for every model. | True on AgentDojo, false on proof-backed data: 32 of 32 cells positive. |
Broad-comparison cascade block-only F1 0.75173. |
That figure was scored against an all-allow stand-in for the rule tier. With the real
rule engine the cascade as it runs today scores 0.73773; escalating a rule
confirm recovers exactly 0.75173, which is why that figure is
the recommendation’s. |
Production-weighted cascade block-only F1 0.12854. |
0.08215 with the real rule engine; the stand-in figure inverts the finding. |
| Hosted Jev’s cascade figures excluded as scored on the stand-in tier. | False: its tier file is byte-identical to the real rule-engine predictions
(9d0df1e6b00f…). The exclusion was inferred from OpenJev’s file in
the same directory. |
| Escalate-on-confirm restores 0.75173 in all orderings. | In one ordering; the others reach 0.704385, 0.495208 and 0.465798. |
| DiffusionGemma leads on any-intervention F1. | That was its Q3 cell ranked against other models’ Q2 cells. At the shared format it scores 0.26792 block-only and 0.18527 any-intervention. |
| A leaderboard ranked at each model’s shipped threshold. | The ranking is now at one shared false-block budget; shipped F1 is a column. |
| The adjudicator allows 89.45% of the short-circuited class. | 89.45% is the modal disagreement class (n = 1,337); the short-circuit-discard class (n = 1,346) is 89.08%. |
| 341,828 normalised rows had no user request. | That is an adapter docstring’s snapshot over 12 corpora; the catalogs record 413,180 over 18. |
| 587 new grade-A families against a prior 16. | 405 new proof-verified families (802 cases before de-duplication) against 333 grade-A rows inventory-wide, 16 in the development split. |
| All 21 fault classes produce allow. | 28 classes were injected; 24 end in allow, 3 of them with no error. |
| ~700,000 provider requests at a 0.0065% error rate. | 1,808,873 requests across 168 manifests; 33 errors, 0.00182%. |
| A $0.0491 per million planning price for the hosted API. | Every priced manifest records $0.042 per million input tokens. |
| 16.67% is a structural prefix-cache ceiling; identical replay hits 89.06%. | 16.67% is exact on a 20-case workload and the same arm scored lower on 200 cases; 89.06% is the 112-token cache unit only, 22.03% at the default 784. |
| Background-subtracted cache shares 25.89% → 44.65% → 62.69% for the prompt-layout arms. | The subtraction is persisted in no artifact; the raw window shares are 11.97%, 18.24% and 33.00%. |
| ~10% of flagged events are unstable across replays. | Per model: 0.00% for OpenJev (0 of 143 flagged events); 17.24% for DiffusionGemma (10 of 58 flagged events); 13.82% for Jev 1.13.0 (21 of 152 flagged events). |
evaluation-only never-train
Nothing derived here is approved for training, synthetic generation, teacher context, distillation or redistribution. The 13 source datasets are public and linked on reproduce; the analysis artifacts behind each number are held privately and are available on request.
Upstream defects reported:
vllm-project/vllm#58015,
wfzyx/von#11.