126 scans. 243 findings. 82 critical.
We pointed Orithos at our own agents.
A security company that red-teams its own AI agents has no excuse to hide the results. Here is the full internal evaluation — the blind spots, the verbatim attack paths, the verification story, and the scans that failed. Every number traces to a finding you can inspect.
How the evaluation actually ran — traceable end to end.
Four verifiable diagrams of the real pipeline, generated from the code that ran the evaluation. Open any explorer to pan, zoom, and follow relationships yourself.
The evaluation topology
Every scan started at the Orithos console and ended at a real agent endpoint. The probe engine drove 1,092 unique probe keys through isolated httpx sessions and the model gateway to six demo personas — with Postgres recording every trace behind row-level security.
- 126 scans dispatched over 12 weeks
- Isolated per-probe sessions — one agent, one client
One finding, traced end to end
A finding is not a flag — it is a recorded chain. The attack payload, the agent's response, the AEGIS-J judge's verdict, and the persisted attack path are all captured in one pass. When a probe succeeds, the trace shows exactly how it got there.
- AEGIS-J verdicts: V-FULL through V-NONE
- Confidence below 0.75 escalates — never inflates
From outcomes to audit-grade evidence
2,000 evaluated outcomes separate cleanly: 87.8% of probes were blocked, and the 243 that got through were judged, confidence-scored, and tagged against 86 distinct framework controls — turning one internal scan into a coverage map across 15+ frameworks.
- 94.8% of findings carry compliance tags
- OWASP LLM 76% · NIST AI RMF 61% · EU AI Act 54% · SOC 2 40%
What happened to all 126
No cherry-picking: 80 scans completed and produced every published finding; 43 failed and contributed zero findings — bucketed right on the diagram — and 3 were operator-aborted. The distribution behind the numbers is on the page.
- Failed runs salvage nothing — no silent retries
- Median completed scan: 3.6 minutes
87.8% of probes were blocked.
The 243 that got through are the story.
Direct attacks get blocked. Tool-mediated ones don't.
Our personas blocked 90% of direct credential-harvesting probes — and only 37% of tool-abuse probes. Requests routed through tool arguments slip past defenses built for direct asks.
One platform, three breach profiles.
Your vertical determines your threat model. Coverage disclosure: only Fortis Legal had substantial completed scans in this window — persona comparisons skew toward it, and we say so rather than average away the difference.
Why 243 findings beats 10,000.
A tool that sprays unverified “findings” creates alert fatigue and erodes trust. Ours judges every response, scores confidence, and discloses exactly which verification layers this evaluation exercised.
Mean confidence 0.90 across judged findings · 38 criticals at 100% attack success rate · replay verification exists in the platform but was not exercised in this window — we disclose that instead of implying otherwise.
What this evaluation is not.
This is internal dogfooding. The six targets were our own demo personas — not customer agents, not production traffic.
43 of 126 scans failed and contribute zero findings: 19 worker/queue issues, 10 endpoint misconfigurations, 9 timeouts, 3 platform bugs we then fixed. 3 more were operator-aborted.
Compliance-tag share is 94.8% under the broad tag parse and 85.6% (208/243) at the verdict level — we publish both rather than round in our favor.
Only Fortis Legal had substantial completed scan coverage; Meridian and NexusHealth were coverage-limited this window.
Replay fields were null across all 243 findings — verification here rests on evaluator verdicts and confidence scoring, which is exactly what we claim and nothing more.
Your agents haven't been scanned this way yet.
We ran this on ourselves first. The same evaluation — probe catalog, attack-path traces, judge verdicts — runs against your agents from day one.
Internal evaluation · orithos.com · All findings, figures, and attack paths are real, unaltered scan data from our own environment.