Campaigns
Real campaigns, really executed. Open any one to read the proving transcripts, the detector signals, the control coverage and the evidence ledger — which is re-verified when the page renders.
These run against our own adversarial fixture, not a customer’s system — an agent we wrote to be vulnerable in the ways real deployments are. The engine is given no knowledge of its rules, so every finding below was genuinely discovered. Each agent appears twice: once before remediation, once after.
| Agent | Started | Attempts | Findings | Severity | Coverage | Adversary |
|---|---|---|---|---|---|---|
| ACME Voice Supportv2.0.0 | 2026-10-07 03:31 | 37 | 0 | clean | 87% | deterministic |
| ACME Voice Supportv1.0.0 | 2026-10-07 03:31 | 39 | 4 | 3 crit1 high | 87% | deterministic |
| ACME Supportv2.0.0 | 2026-10-07 03:31 | 40 | 0 | clean | 93% | deterministic |
| ACME Supportv1.4.2 | 2026-10-07 03:31 | 42 | 11 | 5 crit5 high1 med | 100% | deterministic |
Runs marked deterministic used the offline adversary, which follows fixed strategy ladders and does not adapt its wording to the agent’s replies. Those results are a floor on exploitability, not a ceiling. Set OPENROUTER_API_KEY to enable the adaptive adversary and the rubric judge.