Orithos Guard is live as Arx. It sits inline on agent tool calls, returns allow / escalate / block before the action runs, and compiles your rules into typed gates without a model in the compile path. This post is the launch — the capability, the routing, the cost of asking, and an honest ledger of what is not built yet.
The gap we kept hitting
Most of the agent-security conversation is about the model: whether it can be tricked, whether it will refuse, whether a filter catches a bad string in a prompt. Those are real questions. They were not the ones blocking the teams we talked to.
The blocking question was narrower. A tool call is about to happen. It is legitimate in one context and destructive in another, and nothing in the request path knows the difference. A prompt filter sees text. A sandbox sees a system call. Neither one knows your policy, and neither one is asked for an opinion before the action runs.
The moment that needs a decision is the second before the tool executes — and almost nothing is standing there.
So we built for that second. Not a better filter, and not another dashboard that explains an incident afterwards.
What Arx is
Arx is a pre-execution gate. Your runtime calls it immediately before a tool executes and gets back a verdict: allow, escalate or block. The gate loads the ruleset in force for your org, makes one batched judging call, routes the answers through fixed thresholds, and writes the decision down.
It is not a sandbox and not a scanner bolted onto a chat window. It is a control in the path of the action — which is also exactly why it cannot cover everything. If a tool call never reaches the gate, the gate has no opinion about it. That wiring is the integration work, and we would rather say so here than let you find it in production.
Rules you can read, compiled without a model
A rule is one plain-English sentence plus what kind of rule it is. You write the sentence and pick the type; the guard compiles it into typed questions with fixed routing. There is no model anywhere in the compile — it is arithmetic.
That determinism matters for a boring reason and a serious one. The boring reason: it is fast and reproducible, with no per-rule inference bill and no drift between two compilations of the same rule. The artefact is hash-pinned, so you can tell whether the policy in force changed. The serious reason: a rule you can read is a rule you can argue with. When a verdict looks wrong, you point at the sentence that is wrong instead of at a model's behaviour.
How a verdict is made
The judge returns a probability, not a decision. The routing turns that probability into an action, and the thresholds depend on what kind of rule you wrote — a prohibition and a requirement do not fail the same way.
One thing worth being precise about: escalate does not mean a human somewhere reviews it. There is no review queue in Arx. It means the decision record carries the call, the rule and the confidence, so your own workflow can pick it up. We would rather under-promise that than staff a promise we have not built.
Every rule also carries a fail mode, and the default is closed: if the judge cannot be reached, the call is not waved through. An integrator can deliberately set a rule to open, and then an outage produces an allow that carries a degraded reason on the record — an informed choice, never a silent pass.
What one call carries
Judging is batched. The gate does not call the judge once per rule; it groups the gates and asks once, inside a budget it enforces itself.
| Value | |
|---|---|
| Gates judged per call | up to 25 |
| Per-call token budget | 28,000 |
| Two-gate ruleset (the example above) | 187 tokens — under 1% of budget |
Batching is not only about cost. One call with 25 gates beats 25 calls: fewer round trips, less variance, and one place for the answer to come back. The constraint it creates is real, though — a ruleset whose gates exceed the batch ceiling takes more than one call, and that is a latency you are choosing.
What we measured, and what it does not cover
We have a calibration set of labelled agent turns and we have replayed it. Here is what came back, including the parts that are not flattering.
| Metric | Result | What it is |
|---|---|---|
| Agreement with a second judge | 81.9% | Shadow agreement, 856-turn adaptive-conversation subset |
| Breach recall | 0.76 | On labelled turns in that subset |
| Precision | 0.44 | On the same set |
| Crafted adversarial breaches detected | 18/18 | Fixed hand-built corpus |
| Crafted malicious MCP calls flagged | 7/7 | Same corpus, tool-call shaped |
Read those as what they are, because the labels are load-bearing:
- 81.9% is agreement with a second judge, not ground truth and not a detection rate in the wild. It does not cover the MCP tool-call surface, which needs its own soak before we would quote a number for it.
- 0.76 recall means roughly a quarter of the labelled breaches on that set were missed. Precision on the same set is 0.44. Those are the numbers we are working on, and the ones we would want you to ask us about.
- 18/18 and 7/7 are crafted cases — a fixed corpus, not a population. Passing a corpus you helped build is a floor, not a ceiling.
The routing split above — 83.0% allowed (743 turns), 10.4% escalated (93), 6.6% blocked (59) — is measured on the calibration set of 895 turns, not on customer traffic. We would rather put the precision figure on the launch page than have you find it in an audit.
Ships today — and what does not
A launch post that only lists the good column is a trap for the reader. So here is the whole ledger.
| Capability | Today |
|---|---|
| Gate | Hosted relay, or the same image self-hosted in your own VPC. allow / escalate / block with per-rule fail modes. |
| Rules | Author rules in the console or over the API, preview the compiled gates before saving, or install a conservative starter pack and edit it. |
| Judge key | Ours by default, or bring your own. The org's key is encrypted at rest and resolved before the call. |
| Evidence | Decision log with filters and CSV export. Raw payloads never travel — the record keeps a SHA-256 digest of the state. |
| Retention | Decision records by plan: 7, 90 or 365 days, enforced by purge. |
| Roles | Not yet. One user per org — no owner/admin split, no invites. |
| Review queue | Not built. Escalations are recorded for your workflow; there is no staffed human queue. |
| Environments | Not a per-plan concept — provisioning policy, not a feature you configure. |
Publishing the right column is not modesty. Each row there is something you would otherwise discover during procurement, and we would rather you spent that meeting on your own threat model.
Start
Sign up and you can have a ruleset in force in a few minutes: install a starter pack, look at what it compiled to, and point one runtime at the gate.
- Create an org and mint a token with the
guard:evaluatescope. - Load a ruleset — a starter pack, or your own sentences.
- Call the gate from your runtime, immediately before a tool executes.
- Read the decisions. Tighten the rules that were wrong.
Two positions we are not going to hedge, because they are the reason this exists. Security should not be gated on company size — a two-person team shipping an agent deserves the same control as a bank. And no model gets a pass, ours included: the judge is a component, it is fallible, and the numbers above are how we keep score on it.
It is live today, at console.orithos.com. The contract is at arx.orithos.com/docs. If it gets something wrong, the decision record will tell you which rule did it — and that is the whole design.