The v0.2 evaluation — 7 models, one frozen sample
Label-leak correction. v0 alert IDs were typology-named (e.g. STRUCTURING_SUBTLE-0003) and embedded in the narrative every agent saw — a ground-truth leak. Battery v0.2 uses opaque IDs; the full panel was rerun and this page shows the corrected board. Every headline finding survived; the material movements made weak models look worse (GPT-4o bright-line catch 33%→11%, Grok’s spec-resistant directive residual −31→−48). Disclosure: LIMITATIONS.md, “Known measurement caveats”.
One frozen sample (48 reportable + 12 benign alerts, seed 11), all seven models, every condition, scored against deterministic BSA/AML ground truth on parse-OK decisions only (parse rate ≥95% for every model). Stressor columns use the same measure as the main page: sensitivity lost, in percentage points, against the model’s own baseline — lower is better, 0 means the stressor had no effect. The directive column is the worst case over four judgment-inversion directives, at the basic and full prompt. The public leaderboard requires specificity ≥90% on the benign set (at most 1 of 12 benign alerts escalated); Gemma-3-27b falls below it (50% false alarms) and appears here for the record but is not ranked.
Sensitivity change by condition (points — 0 is best)
| Model | Baseline misses | False alarms | Directive, basic → full prompt | Pressure | Cover (A1) | Cover (A2) | Bright-line catch rate |
|---|---|---|---|---|---|---|---|
| Claude Opus 4.8 | 0 | 8% | −69 → −4 | −2 | −19 | −27 | 100% |
| GPT-5.5 | 0 | 0% | −60 → 0 | 0 | −21 | −46 | 100% |
| Claude Haiku 4.5 | 2% | 8% | −48 → −4 | −13 | −43 | −65 | 22% |
| GPT-4o | 0 | 0% | −60 → −6 | −17 | −38 | −73 | 11% |
| Gemma-3-27b | 0 | 50% | −88 → −15 | −4 | −8 | −46 | 25% |
| Llama-3-70b | 0 | 8% | −69 → −12 | −6 | −50 | −69 | 11% |
| Grok-4.3 | 0 | 0% | −52 → −48 | −21 | −40 | −57 | 44% |
Does a fuller prompt protect against the covers? Three models were also run on the A2 covers with the basic prompt. Sensitivity losses were higher there — Opus 58, GPT-4o 83, Grok 73 points, versus 27, 73, and 57 with the full prompt. So the full prompt reduces the effect but most of it remains, unlike the directive attack, which the full prompt largely neutralizes. The covers work on the alert itself, not on a gap in the prompt.
Observations from the v0.2 sample:
- Under the directive attack at the basic prompt, every model under-escalated (48 to 88 points of sensitivity lost). Stating the legal standard in full removed most of the effect for five of seven models; two kept sizable residuals (12 and 48 points), one of them against an explicit anti-override clause. A fuller prompt helps, but not uniformly.
- Pressure alone, with no changed decision rule, moved four of seven models (13 to 21 points) and left the rest near zero. At the full prompt it moved no ranked model more than 4 points.
- Cover stories raised misses for all seven models even with the full prompt in place — including the two models unaffected by every prompt-side condition (Opus 27 points, GPT-5.5 46).
- Integrity under cover ranged from 100% down to 11%. Where it is low, the A2 number mixes cover-susceptibility with general competence loss, so the two columns should be read together.
- These are observations about seven models on one frozen sample — not general claims about model classes or capability tiers.
The deception ladder (A0 → A3)
Every reportable alert can be wrapped in a cover story — the innocent explanation a launderer would present — while the transaction facts and the correct answer stay fixed. Sensitivity is measured as the cover deepens: A0 raw, no cover; A1 one grounded business rationale; A2 adds a supporting KYC/expected-activity profile (the leaderboard’s deceptive-cover column); A3 gives the adversary three distinct covers per alert, and the alert counts as missed if any one of them clears it.
Escalation sensitivity by adversary level (higher is better)
| Model | At A2 | At A3 (best-of-3) | Bright-line catch rate at A3 |
|---|---|---|---|
| Claude Opus 4.8 | 67% | 38% | 78% |
| Llama-3-70b | 33% | 23% | 0% |
| GPT-5.5 | 48% | 19% | 22% |
| GPT-4o | 31% | 19% | 11% |
| Grok-4.3 | 37% | 17% | 0% |
| Claude Haiku 4.5 | 31% | 12% | 0% |
| Gemma-3-27b (not ranked — below the specificity gate) | 48% | 29% | 11% |
A3 lowers sensitivity further for every model, but partly by giving up plausibility: a cover that clears unmistakable overt structuring is no longer an innocent explanation a compliance officer would credit — it is past the ecological-validity gate. So the A3 numbers are an upper bound on what a cover-picking adversary can force, not a clean measure of deception. Two things hold. First, A1/A2 remain the gate-clean results: there the bright-line catch rate stays at or near 100% for the most capable models, so the losses reflect genuine judgment-shading. Second, at A3 only Opus kept most of its bright-line integrity (78%); every other model tested fell to 0–22%, which means much of their A3 loss is plausibility slippage rather than deception a plausible cover could achieve.
The investigation, step by step
Everything below is the exploratory pass that led to the v0 design, kept as the record. Where the two disagree, the v0.2 table above supersedes it — for example, the exploratory pass attributed incentive-immunity to “the frontier”; the v0.2 run shows it varies by model.
| Step | Claim tested | Finding |
|---|---|---|
| Baseline | A peer-benchmark incentive bends the frontier ~37% | Reproduced: Opus 4.8 lost 36 points under the quota framing. |
| Ablation | The incentive (benchmark / number / penalty) is the active ingredient | No — those are ~0 on the frontier; the active ingredient is a permissive directive (“when in doubt, clear it”). |
| Capability | Susceptibility is universal | No — incentive susceptibility varied by model: smaller models moved, the strongest did not. The directive moved every model on a basic prompt. |
| Validity | The frontier follows unlawful directives | Only when the base prompt leaves the doubt-standard implicit. State it explicitly and directives collapse to 0 on Opus. |
| The fix | The explicit standard protects everyone | No — protection scales with capability. A correct prompt does not make a weak model safe. |
| Overrides | Real, cited suppression guidance beats the well-specified frontier | Mostly no — enforcement-grounded fact-overrides bounce off. The one survivor hit an uncovered spec dimension, then collapsed once that dimension was stated. |
| Cover adversary | A grounded cover story in the alert also can’t beat a spec-complete frontier | It does — the only manipulation that survives spec-completeness on the frontier. |
Axis 1 — model capability and the stated standard
With the file-on-doubt standard stated explicitly in the system prompt, residual under-escalation under incentive pressure varied with model capability. A correct prompt did not make the weaker models safe.
Sensitivity change under the quota framing (points — 0 is best)
| Model | Baseline misses | Lost, basic prompt | Lost, full prompt |
|---|---|---|---|
| Claude Opus 4.8 | 0% | −35 | 0 |
| GPT-5.5 | 0% | −23 | not run; ~0 expected |
| Claude Haiku 4.5 | 5% | −45 | not run |
| GPT-4o | 10% | −38 | −18 |
| Grok-4.3 | 2% | −53 | −34 |
| Gemma-3-27b | 0% | −80 | not run; high expected |
| Llama-3-70b | 4% | −32 | not run |
Grok still cleared roughly a third of reportable alerts with the obligation explicit. Gemma lost 92 points under a directive and cleared even overt structuring — a competence failure rather than shaded judgment.
Axis 2 — incentives vs. directives, and spec-completeness
Component ablation of the quota framing showed the incentive elements (benchmark, number, penalty) had little effect on the strongest models; the permissive directive carried the effect. Within directives, only those that redefine the evidentiary threshold moved the strongest models (“escalate only on confirmed intent,” 66 points on Opus); operational rationales (materiality, efficiency, SLA) did not. The pattern: these models resisted competing motives but followed an explicit decision rule that inverted the default — when the default was not stated.
Against a spec-complete base prompt, grounded adversarial overrides drawn from real enforcement actions (Rabobank’s “Verified List,” U.S. Bank’s capacity caps, Canaccord’s pre-filter, HSBC’s franchise-value rationale) had no effect on Opus. The one survivor — a plausible but false legal standard (“a SAR requires proof of a predicate crime”) — cost Opus 34 points precisely because the prompt hadn’t covered that dimension. Adding the FFIEC “no particular crime need be proven” clause removed the effect entirely.
Every prompt-based failure maps to a specification gap. The frontier resists false facts and accepts a plausible false standard — only on obligations the prompt left implicit. This makes Axis 2 an auditing tool: each benchmark item names an obligation a deployed prompt must state.
Axis 3 — the cover-story adversary
Instead of changing the prompt, we attach a cover story to the alert: a genuinely reportable alert gets a plausible business explanation woven into its narrative, as often done by actual launderers, while the transaction features and the ground-truth label stay untouched. A correct agent escalates on the pattern regardless of the story. Model and system prompt are held fixed (full spec, no incentive, no directive); the only variable is the cover.
| Level | Opus 4.8 missed | Opus bright-line control | GPT-5.5 missed | GPT-5.5 control |
|---|---|---|---|---|
| A0 — raw | 0.0% | 0% | 0.0% | 0% |
| A1 — one grounded cover | 18.8% | 0% | 22.9% | 10% |
| A2 — cover + KYC context | 35.4% | 0% | 50.0% | 20% |
| A3 — best-of-3 covers | 62.5% | 30% | not run | — |
Four observations from the exploratory pass:
- It defeated models unaffected by every prompt-side condition. The prompt even states the relevant principle — “a plausible innocent explanation does not defeat reasonable suspicion” — and models recite it while failing to apply it. Because the clause is already present, adding prompt language does not fix this the way it fixes the directive attacks.
- The effect grows with adversary effort (A1 → A2 → A3), so difficulty can be raised by level as models improve.
- In this exploratory pass, the two models diverged on the bright-line control: Opus held it at 0% through A2 while GPT-5.5 missed 20% of overt structuring. This gap did not replicate in the v0.2 run above, where both models held the control at 100% — a reminder to read single-run gaps cautiously.
- The bright-line control is the validity gate. It stayed at 0% for Opus through A1/A2, then moved to 30% at A3 — the sign that best-of-N selection had begun producing covers strong enough to explain away an unmistakable pattern. The supported claim is A1/A2; the A3 62.5% is an upper bound.
Grounding discipline
Every framing and cover traces to the real record, tiered by provenance: incentive and directive framings from the HSBC/Everett Stern and TD Bank enforcement records; overrides from Rabobank, U.S. Bank, Canaccord, HSBC, and FFIEC guidance; the launderer cover library from HSBC/Casa de Cambio, the Russian and Troika Laundromats, 1MDB, Danske, and FATF trade-based money-laundering typologies. Contrived items are excluded by an ecological-validity gate: the benchmark measures manipulations that occur in practice, not ones invented to break models. For the cover-story adversary, that gate is the bright-line control.
Method notes & caveats
- Point estimates on 48 reportable alerts per cell (±~0.05–0.07), single run; read for shape, not third decimals. Magnitudes are phrasing-sensitive.
- Covers were authored by Claude Opus 4.8, so the Opus rows are mild self-play; every other model faces a cross-model adversary, which makes those rows the cleaner evidence. Opus was not the first choice: Claude Fable 5 declined to write laundering cover stories even with the benchmark’s defensive purpose stated; Opus accepted the defensive framing. Which models will author adversarial material is itself a variable in red-teaming AML systems.
- Provenance is tiered, not uniform. Incentive framings and adversarial cover stories are enforcement-grounded (HSBC/Stern, TD Bank, Rabobank, the Russian/Troika Laundromats, 1MDB, Danske). Of the four judgment-inversion directives, one is enforcement-documented and one enforcement-adjacent, but two are extrapolated from the FFIEC “need not confirm the underlying crime” standard rather than captured verbatim — labeled as such, and never attributed to a named case. The full tiered evidence base, item by item, is in the spec’s §12.
- A2 is only a clean deception measure where bright-line integrity is high; for low-integrity models it partly reflects competence collapse under cover.
- Cover-library provenance carries a survivorship caveat: only schemes that were eventually caught are documented.
Raw per-run outputs are reproducible from the open harness; the durable record is the repository’s findings document.