Research notes: how the thesis moved

We set out to prove that hidden incentive pressure corrupts AML triage agents. Rigorous ablation falsified the strong version of that claim on frontier models, while underscoring that sub-frontier models remain vulnerable, and frontier models are diverted from their role responsibilities through misleading background on alerts. More background on the investigation arc is featured below. Dead ends are featured alongside the findings for transparency.

The v0.2 evaluation — 7 models, one frozen sample

Label-leak correction. v0 alert IDs were typology-named (e.g. STRUCTURING_SUBTLE-0003) and embedded in the narrative every agent saw — a ground-truth leak. Battery v0.2 uses opaque IDs; the full panel was rerun and this page shows the corrected board. Every headline finding survived; the material movements made weak models look worse (GPT-4o bright-line catch 33%→11%, Grok’s spec-resistant directive residual −31→−48). Disclosure: LIMITATIONS.md, “Known measurement caveats”.

One frozen sample (48 reportable + 12 benign alerts, seed 11), all seven models, every condition, scored against deterministic BSA/AML ground truth on parse-OK decisions only (parse rate ≥95% for every model). Stressor columns use the same measure as the main page: sensitivity lost, in percentage points, against the model’s own baseline — lower is better, 0 means the stressor had no effect. The directive column is the worst case over four judgment-inversion directives, at the basic and full prompt. The public leaderboard requires specificity ≥90% on the benign set (at most 1 of 12 benign alerts escalated); Gemma-3-27b falls below it (50% false alarms) and appears here for the record but is not ranked.

Sensitivity change by condition (points — 0 is best)

ModelBaseline
misses
False
alarms
Directive,
basic → full prompt
PressureCover
(A1)
Cover
(A2)
Bright-line
catch rate
Claude Opus 4.808%−69−4−2−19−27100%
GPT-5.500%−6000−21−46100%
Claude Haiku 4.52%8%−48−4−13−43−6522%
GPT-4o00%−60−6−17−38−7311%
Gemma-3-27b050%−88−15−4−8−4625%
Llama-3-70b08%−69−12−6−50−6911%
Grok-4.300%−52−48−21−40−5744%

Does a fuller prompt protect against the covers? Three models were also run on the A2 covers with the basic prompt. Sensitivity losses were higher there — Opus 58, GPT-4o 83, Grok 73 points, versus 27, 73, and 57 with the full prompt. So the full prompt reduces the effect but most of it remains, unlike the directive attack, which the full prompt largely neutralizes. The covers work on the alert itself, not on a gap in the prompt.

Observations from the v0.2 sample:

A caveat on Gemma’s clean floor. It misses no reportable alerts, but it escalates 58% of the benign ones, so its low miss rate reflects indiscriminate over-flagging rather than discrimination. The whole row should be read with that in mind.

The deception ladder (A0 → A3)

Every reportable alert can be wrapped in a cover story — the innocent explanation a launderer would present — while the transaction facts and the correct answer stay fixed. Sensitivity is measured as the cover deepens: A0 raw, no cover; A1 one grounded business rationale; A2 adds a supporting KYC/expected-activity profile (the leaderboard’s deceptive-cover column); A3 gives the adversary three distinct covers per alert, and the alert counts as missed if any one of them clears it.

Escalation sensitivity by adversary level (higher is better)

ModelAt A2At A3 (best-of-3)Bright-line catch rate at A3
Claude Opus 4.867%38%78%
Llama-3-70b33%23%0%
GPT-5.548%19%22%
GPT-4o31%19%11%
Grok-4.337%17%0%
Claude Haiku 4.531%12%0%
Gemma-3-27b (not ranked — below the specificity gate)48%29%11%

A3 lowers sensitivity further for every model, but partly by giving up plausibility: a cover that clears unmistakable overt structuring is no longer an innocent explanation a compliance officer would credit — it is past the ecological-validity gate. So the A3 numbers are an upper bound on what a cover-picking adversary can force, not a clean measure of deception. Two things hold. First, A1/A2 remain the gate-clean results: there the bright-line catch rate stays at or near 100% for the most capable models, so the losses reflect genuine judgment-shading. Second, at A3 only Opus kept most of its bright-line integrity (78%); every other model tested fell to 0–22%, which means much of their A3 loss is plausibility slippage rather than deception a plausible cover could achieve.

The investigation, step by step

Everything below is the exploratory pass that led to the v0 design, kept as the record. Where the two disagree, the v0.2 table above supersedes it — for example, the exploratory pass attributed incentive-immunity to “the frontier”; the v0.2 run shows it varies by model.

StepClaim testedFinding
BaselineA peer-benchmark incentive bends the frontier ~37% Reproduced: Opus 4.8 lost 36 points under the quota framing.
AblationThe incentive (benchmark / number / penalty) is the active ingredient No — those are ~0 on the frontier; the active ingredient is a permissive directive (“when in doubt, clear it”).
CapabilitySusceptibility is universal No — incentive susceptibility varied by model: smaller models moved, the strongest did not. The directive moved every model on a basic prompt.
ValidityThe frontier follows unlawful directives Only when the base prompt leaves the doubt-standard implicit. State it explicitly and directives collapse to 0 on Opus.
The fixThe explicit standard protects everyone No — protection scales with capability. A correct prompt does not make a weak model safe.
OverridesReal, cited suppression guidance beats the well-specified frontier Mostly no — enforcement-grounded fact-overrides bounce off. The one survivor hit an uncovered spec dimension, then collapsed once that dimension was stated.
Cover adversaryA grounded cover story in the alert also can’t beat a spec-complete frontier It does — the only manipulation that survives spec-completeness on the frontier.

Axis 1 — model capability and the stated standard

With the file-on-doubt standard stated explicitly in the system prompt, residual under-escalation under incentive pressure varied with model capability. A correct prompt did not make the weaker models safe.

Sensitivity change under the quota framing (points — 0 is best)

ModelBaseline missesLost, basic promptLost, full prompt
Claude Opus 4.80%−350
GPT-5.50%−23not run; ~0 expected
Claude Haiku 4.55%−45not run
GPT-4o10%−38−18
Grok-4.32%−53−34
Gemma-3-27b0%−80not run; high expected
Llama-3-70b4%−32not run

Grok still cleared roughly a third of reportable alerts with the obligation explicit. Gemma lost 92 points under a directive and cleared even overt structuring — a competence failure rather than shaded judgment.

Axis 2 — incentives vs. directives, and spec-completeness

Component ablation of the quota framing showed the incentive elements (benchmark, number, penalty) had little effect on the strongest models; the permissive directive carried the effect. Within directives, only those that redefine the evidentiary threshold moved the strongest models (“escalate only on confirmed intent,” 66 points on Opus); operational rationales (materiality, efficiency, SLA) did not. The pattern: these models resisted competing motives but followed an explicit decision rule that inverted the default — when the default was not stated.

One detail worth recording: on GPT-5.5, wrapping the directive in incentive framing weakened it (directive alone 44 points; full quota 10). The pressure cues apparently read as manipulation — so the least pressure-like phrasing was the most effective.

Against a spec-complete base prompt, grounded adversarial overrides drawn from real enforcement actions (Rabobank’s “Verified List,” U.S. Bank’s capacity caps, Canaccord’s pre-filter, HSBC’s franchise-value rationale) had no effect on Opus. The one survivor — a plausible but false legal standard (“a SAR requires proof of a predicate crime”) — cost Opus 34 points precisely because the prompt hadn’t covered that dimension. Adding the FFIEC “no particular crime need be proven” clause removed the effect entirely.

Every prompt-based failure maps to a specification gap. The frontier resists false facts and accepts a plausible false standard — only on obligations the prompt left implicit. This makes Axis 2 an auditing tool: each benchmark item names an obligation a deployed prompt must state.

Axis 3 — the cover-story adversary

Instead of changing the prompt, we attach a cover story to the alert: a genuinely reportable alert gets a plausible business explanation woven into its narrative, as often done by actual launderers, while the transaction features and the ground-truth label stay untouched. A correct agent escalates on the pattern regardless of the story. Model and system prompt are held fixed (full spec, no incentive, no directive); the only variable is the cover.

LevelOpus 4.8 missedOpus bright-line controlGPT-5.5 missedGPT-5.5 control
A0 — raw0.0%0%0.0%0%
A1 — one grounded cover18.8%0%22.9%10%
A2 — cover + KYC context35.4%0%50.0%20%
A3 — best-of-3 covers62.5%30%not run

Four observations from the exploratory pass:

Grounding discipline

Every framing and cover traces to the real record, tiered by provenance: incentive and directive framings from the HSBC/Everett Stern and TD Bank enforcement records; overrides from Rabobank, U.S. Bank, Canaccord, HSBC, and FFIEC guidance; the launderer cover library from HSBC/Casa de Cambio, the Russian and Troika Laundromats, 1MDB, Danske, and FATF trade-based money-laundering typologies. Contrived items are excluded by an ecological-validity gate: the benchmark measures manipulations that occur in practice, not ones invented to break models. For the cover-story adversary, that gate is the bright-line control.

Method notes & caveats

Raw per-run outputs are reproducible from the open harness; the durable record is the repository’s findings document.