the adversarial audit
It isn't one audit. It's a discipline.
Anyone can run a security pass once and screenshot the green. The harder thing is to keep re-attacking your own system as it grows, and get the same answer every time. The latest round: a 50-agent audit, its fixes shipped, then a separate 14-lens adversarial re-audit where every finding had to survive a skeptic before it counted. What holds, holds.
what the latest audit covered
Scope, method and limits first. Agent counts after.
From the latest core audit, dated 25 September 2026. The number of agents says how busy an audit was, not how good it was, so here is what this one actually did.
- Code
- The core package, 162 files and 17,362 lines, at one fixed commit, plus the gateway code needed to tell whether each path is live.
- Access
- The full source, and a fresh database built from the migrations. Not the production database; production-only settings were marked unverified and checked by me afterwards.
- Method
- Four reviewers each read one layer in full. A lead re-read every finding against the file, and ran every timing or pattern claim, before it counted.
- Severity
- HIGH means a missed crisis or a denial of service on a live path, a cross-tenant leak, or money lost. Medium means a guard that fails open, a credit created from nothing, or data silently dropped. Every finding is also marked live or latent: reachable in production today, or not.
- Disagreement
- A finding the lead could not reproduce was dropped or marked unverified, not averaged in.
- Result
- 2 HIGH and 13 medium findings. All fixed the same day, each with a test that failed on the old code, and the full suite re-run in CI.
- Not covered
- The fidelity of about 7,000 lines of translated safety strings, the production database, the web server configuration, and 17 low findings still open.
Self-run by AI reviewers I directed, not independent third-party assurance.
the cadence
Re-attacked, round after round. Every cross-tenant finding below was found by an audit, not reported by a customer, and fixed.
✓ crown jewels held every round, isolation · money · governance · SSRF · token domains
three of the real ones
HIGH Cross-tenant conversation leak
Chat memory was keyed by the sender, not the tenant, so a guest who messaged two businesses through the platform could have one's conversation surface inside the other's prompt. Fixed: history is keyed by business at the storage layer, with a test that fails if the old key comes back. It was found by the audit, not reported by a customer.
HIGH A provider draining a caller
In the cross-tenant tool economy, a provider could raise the price of a granted tool after a caller began using it and drain their balance. Fixed: the caller consents to a price per call, fail-closed, no consent, no charge, enforced before a single credit moves.
HIGH Crisis behind the meter
The billing check ran before the crisis check, so a distressed customer messaging an out-of-credits business got "unavailable" instead of a crisis line. Fixed: crisis is served free, ahead of every ceiling and charge, on every customer surface.
The lenses span isolation · auth · injection · money · governance · booking · channels · the autonomous resident · data-rights (GDPR/PECR) · the site generator · frontend · ops. Verdict, every round: zero critical standing; tenant isolation and the XSS surface came back clean; no live auth bypass: every confirmed edge fixed and redeployed.
the honest numbers
Every number, with the asterisk already attached.
Most builders pad. I'd rather you trust the parts that are real than be impressed by parts that aren't.
The system is honest because the person who designed it is. I made sure of that from the start, built in, not bolted on.
the invitation
The numbers hold. Asterisks included.
If you want a builder who attacks his own work this hard before anyone asks, and keeps doing it as the system grows, let's talk.
Ask about the audits: what was found, what is fixed, what is still open.
- What is still open?
- What was the worst thing found?
- Who ran these audits?