AgentForge: Clinical Co-Pilot resilience

Continuous adversarial testing, over time. Read-only; agent config and campaign controls live on the Configuration page.
public view  Log in →
loading

Resilience over SUT versions (defended %)

loading…
version timeline (how results changed as the target was improved)
loading…

Test catalog

filter:
Every attack angle the Red Team can run. Click a test that has run to see its actual messages and verdicts across versions. Greyed angles are defined but have not run yet.
loading…

Run history

Each campaign execution. Click a run to see the models it used and (once you log in) the full chat transcript.
loading…

Findings & regression outcomes

loading…

Published reports

loading…
What the terms mean
Attempt
One adversarial test case sent to the SUT (possibly several conversational turns), then scored by the Judge. It is the denominator: 30 attempts / 0 findings means the target defended all 30.
Finding
An attempt the Judge ruled a genuine weakness (a fail or partial), recorded as a vulnerability. Findings are a subset of attempts. Zero findings means the SUT defended every attempt.
Not judged
An attempt that was sent to the SUT but has no verdict: the Judge returned a typed error (a timeout, or a rate-limit / credit exhaustion during a large run). The platform records the error and never fabricates a pass, so the attempt stays unscored rather than counting as defended.
Published report
A confirmed vulnerability written up as a report. A critical finding stays pending until an operator approves it, then it is published here. Click one to read the reproduction and detail.
Regression outcome
When the SUT version changes, each open finding is re-tested. "Resolved" means the fix held; "reappeared" means it came back. This is how a fix is verified over time.
Test angle
A specific probe (category · subcategory, e.g. prompt-injection · role-abandon). The catalog lists every angle the Red Team can run; one only charts a trend once it has actually run. The Red Team writes a fresh, varied attack each time it runs an angle (breadth over replaying one identical string), so the message differs run to run by design.
SUT / "System Under Test"
The deployed Clinical Co-Pilot we attack. Each run is stamped with the SUT version it hit; live means the currently-deployed target, not pinned to a specific build.
Defended %
Share of a version's attempts the SUT passed (higher is more resilient).