Research & Articles
What we find, and how we think about it.
Probe methodology, published findings, and platform writing from the team building Orithos.
5stages, one call
25gates per batch
83.0%allowed, calibration set
Arx is live: a gate for agent tool calls
We shipped a pre-execution gate for agent tool calls: rules you can read, compiled without a model, and a decision record for every verdict. What ships today, what it costs to ask, and the numbers that do not flatter us.
Latest12 posts
Model-selected
LLM-selected
Δ within noise at this sampleNull result — publishedMETHODOLOGY
The A/B we couldn't win: our decision model picks attack techniques now — here's what the data actually said
We gave our non-generative decision model a bigger job: choosing which attack technique each adversarial turn uses, with the LLM writing the payload. Then we ran the experiment that could prove it beats the LLM at that job. It didn't — not at this sample. We're publishing the null result.
81.9%61.2%
GUARDRAILS
Judging the judges: what 895 adversarial turns taught us about guardrail models
We replayed 856 identical adversarial turns through two candidate guardrail judges. The stronger one agreed with our production judge on 82% of turns; the open-source alternative ranged from 20% to 61%, depending on which serving endpoint answered. The headline number was the least interesting result.
10.4% judge-routed89.6% resolved below the frontier
9.5× cheaper · 0 invariant breaksGUARDRAILS
The exception band: cutting guard cost 9.5× without weakening the gate
Continuous agent guarding dies on unit economics. If every turn reaches a frontier judge, safety gets sampled — not enforced. Here's the architecture that skipped 89.6% of judge calls, and the invariants that kept it honest.
Verifyclaim checked
Attestevidence minted
Storehash-tied record
re-verifies continuouslyGUARDRAILS
Your evidence ages. We made ours re-verify.
Every audit, every security review, every "are we compliant?" conversation has the same hidden defect: the evidence was true when it was minted. This week we shipped the part that keeps it true — or says so when it isn't.
85 contested turns — every one adjudicated by hand
4 real catches1 weak1 no-catch
GUARDRAILS
Every disagreement, adjudicated: what 85 contested turns taught us about trusting a judge
When two judges disagree, most teams tune a threshold and move on. We did something slower: we read all 85 disagreements by hand and adjudicated each one against the record. The result changed how our guard routes decisions — and resolved a question we'd left open.
Input boundarymost volume lands here
Tool callallow / deny
Outputegress filter
Memorywrite guard
Data flowwhere failures stick
Human in the loopescalation gate
GUARDRAILS
Where guardrails actually need to sit: 227 findings mapped to the ACS hook model
Every probe in catalog v1.3 now declares which ACS v0.1 hook surface a guardrail must cover to block it. We mapped our dogfooding findings to those surfaces — the input boundary absorbs most of the volume, but the failures that cross data-flow boundaries are the ones that stick.
126scans run inward
243findings, disclosed
87.8%of probes blocked
CASE STUDIES
We red-teamed our own AI agents: 126 scans, 243 findings, and what it taught us
A quarter of dogfooding turned inward: 126 scans against our own agents, 87.8% of probes blocked — and 243 findings showing exactly where guardrails collapse. Every number computed from the internal dataset; every gap disclosed.
87.8%of probes blocked
2,000evaluated outcomes · 6 agents
243findings — 72% critical / high
BENCHMARKS
State of Agent Security — Q3 2026
The first edition of our quarterly benchmark: 2,000 evaluated probe outcomes across six agents — 87.8% blocked, 243 real findings, 72% of them critical or high. The aggregate data, the method behind it, and the parts that went wrong.
You are here
Aug 2, 2026Enforcement live
Dec 2027High-risk regime
Aug 2028Remainder
EU AI ACT
The EU AI Act is live. Here’s what agent builders actually owe — and in what order.
Enforcement began August 2, 2026. Three obligations apply to agent builders today, the high-risk regime moved to December 2027 — and the sequencing mistake almost everyone makes is doing the paperwork before the testing.