…
loading
Resilience over SUT versions (defended %)
loading…
version timeline (how results changed as the target was improved)
loading…
Every attack angle the Red Team can run. Click a test that has run to see its actual messages and verdicts across versions. Greyed angles are defined but have not run yet.
loading…
Run history
Each campaign execution. Click a run to see the models it used and (once you log in) the full chat transcript.
loading…
Findings & regression outcomes
loading…
Published reports
loading…
What the terms mean
- Attempt
- One adversarial test case sent to the SUT (possibly several conversational turns), then scored by the Judge. It is the denominator: 30 attempts / 0 findings means the target defended all 30.
- Finding
- An attempt the Judge ruled a genuine weakness (a fail or partial), recorded as a vulnerability. Findings are a subset of attempts. Zero findings means the SUT defended every attempt.
- Not judged
- An attempt that was sent to the SUT but has no verdict: the Judge returned a typed error (a timeout, or a rate-limit / credit exhaustion during a large run). The platform records the error and never fabricates a pass, so the attempt stays unscored rather than counting as defended.
- Published report
- A confirmed vulnerability written up as a report. A critical finding stays pending until an operator approves it, then it is published here. Click one to read the reproduction and detail.
- Regression outcome
- When the SUT version changes, each open finding is re-tested. "Resolved" means the fix held; "reappeared" means it came back. This is how a fix is verified over time.
- Test angle
- A specific probe (category · subcategory, e.g. prompt-injection · role-abandon). The catalog lists every angle the Red Team can run; one only charts a trend once it has actually run. The Red Team writes a fresh, varied attack each time it runs an angle (breadth over replaying one identical string), so the message differs run to run by design.
- SUT / "System Under Test"
- The deployed Clinical Co-Pilot we attack. Each run is stamped with the SUT version it hit;
livemeans the currently-deployed target, not pinned to a specific build. - Defended %
- Share of a version's attempts the SUT passed (higher is more resilient).