architecture

What it takes to let an agent act

Wiring a model to a prompt is a demo. The distance between that and something you can put other people's money through is four specific mechanisms, none of them the model.

Most AI products answer questions. The moment one acts, takes a booking, moves money, writes to a record another customer can read: the engineering problem changes completely, and almost none of the new problem is about the model.

PANTHEON is a substrate for running agents that act. Here is what that actually required, in the order the failures would have arrived.

Isolation the application cannot talk its way past

Two unrelated businesses run on one spine. The obvious way to keep their data apart is a WHERE tenant_id = ? in every query, which works until one query forgets.

So the boundary is not in the application. It is Postgres row-level security, forced, on a role that carries NOBYPASSRLS. With no tenant context set, a SELECT returns zero rows, not every row. Fail-closed, because a defence that fails open is worse than none: it produces a passing test and a breach.

You can attack this yourself on /prove.html: a live control read, then the same read across the tenant boundary, against the running system.

A permission an agent cannot grant itself

An agent that can decide it is allowed to do something is not governed. The authorization gate is tiered, and the consequential tier does not resolve to "allowed". it resolves to queued, with an approval ID and a human on the other end.

The kill switch is the same idea, one level up: with it thrown, a tier-3 action is refused before it is even queued.

Metering that cannot overdraft

Credits are decremented atomically, before the work, not after it. The sequence matters: reserve, act, settle. Bill-after-the-fact is how you discover that a retry storm has run up a bill against an account that had four credits.

One consequence I insisted on: a person in crisis reaches help at zero credits. The crisis check runs before the balance check. That ordering is a product decision expressed as a line number.

A judge that does not share the generator's blind spots

The quality layer only spends an expensive LLM judge when the reply is high-stakes: a cheap pre-pass decides. On a low-stakes draft it makes zero judge calls; on a priced booking, one.

And the judge is a different model from the generator. A model reviewing its own output agrees with itself; that is not review, it is a second opinion from the same brain.

That layer is open source and installable:

pip install pantheon-guardrails

What it is not

Scale is designed for and unproven, one production instance, by choice. Prompt injection is mitigated by defence in depth, not solved; nobody has solved it and I will not pretend otherwise. Isolation has been adversarially tested and held, which is strong evidence and not a formal proof.

Those three sentences are on the site because a claim you cannot check is worth nothing, and a limit you volunteer is worth more than a benchmark you chose.


Six components are extracted and public: guardrails, tool-sanitizer, ssrf-guard, rls, ical, credit-ledger, all Apache-2.0, all on PyPI.

← all posts