the adversarial audit

It isn't one audit. It's a discipline.

Anyone can run a security pass once and screenshot the green. The harder thing is to keep re-attacking your own system as it grows, and get the same answer every time. The latest round: a 50-agent audit, its fixes shipped, then a separate 14-lens adversarial re-audit where every finding had to survive a skeptic before it counted. What holds, holds.

what the latest audit covered

Scope, method and limits first. Agent counts after.

From the latest core audit, dated 25 September 2026. The number of agents says how busy an audit was, not how good it was, so here is what this one actually did.

Code
The core package, 162 files and 17,362 lines, at one fixed commit, plus the gateway code needed to tell whether each path is live.
Access
The full source, and a fresh database built from the migrations. Not the production database; production-only settings were marked unverified and checked by me afterwards.
Method
Four reviewers each read one layer in full. A lead re-read every finding against the file, and ran every timing or pattern claim, before it counted.
Severity
HIGH means a missed crisis or a denial of service on a live path, a cross-tenant leak, or money lost. Medium means a guard that fails open, a credit created from nothing, or data silently dropped. Every finding is also marked live or latent: reachable in production today, or not.
Disagreement
A finding the lead could not reproduce was dropped or marked unverified, not averaged in.
Result
2 HIGH and 13 medium findings. All fixed the same day, each with a test that failed on the old code, and the full suite re-run in CI.
Not covered
The fidelity of about 7,000 lines of translated safety strings, the production database, the web server configuration, and 17 low findings still open.

Self-run by AI reviewers I directed, not independent third-party assurance.

the cadence

Re-attacked, round after round. Every cross-tenant finding below was found by an audit, not reported by a customer, and fixed.

latest
14-lens re-audit of the last remediation, residual-bypass + regression hunt per fix, each finding refuted by an independent agent. 2 low edge-cases survived, both fixed; nothing cross-tenant, no money or governance break.
before it
50-agent audit → a minimal 9-PR remediation across the store seams (calendar, channels, the autonomous resident, the tenant economy, GDPR/PECR): all shipped, tested, redeployed.
before that
44-agent audit, 9 findings, 3 high (incl. a cross-tenant conversation leak): all fixed; and a 30-agent extension analysis that surfaced real bugs as a side-effect of grounding.
July 2026
80-agent, 14-facet audit: the earliest full sweep, run with a refutation rule: no finding counted until a separate agent failed to refute it. Three fixes shipped from it, including a quadratic website-importer rewritten to a bounded linear scan. Separate from, and preceding, the 50-agent sweep above.
and earlier
whole-platform and whole-Studio adversarial passes, the flagship, and several before them: a running trail in git.

✓ crown jewels held every round, isolation · money · governance · SSRF · token domains

three of the real ones

HIGH Cross-tenant conversation leak

Chat memory was keyed by the sender, not the tenant, so a guest who messaged two businesses through the platform could have one's conversation surface inside the other's prompt. Fixed: history is keyed by business at the storage layer, with a test that fails if the old key comes back. It was found by the audit, not reported by a customer.

HIGH A provider draining a caller

In the cross-tenant tool economy, a provider could raise the price of a granted tool after a caller began using it and drain their balance. Fixed: the caller consents to a price per call, fail-closed, no consent, no charge, enforced before a single credit moves.

HIGH Crisis behind the meter

The billing check ran before the crisis check, so a distressed customer messaging an out-of-credits business got "unavailable" instead of a crisis line. Fixed: crisis is served free, ahead of every ceiling and charge, on every customer surface.

The lenses span isolation · auth · injection · money · governance · booking · channels · the autonomous resident · data-rights (GDPR/PECR) · the site generator · frontend · ops. Verdict, every round: zero critical standing; tenant isolation and the XSS surface came back clean; no live auth bypass: every confirmed edge fixed and redeployed.

the honest numbers

Every number, with the asterisk already attached.

Most builders pad. I'd rather you trust the parts that are real than be impressed by parts that aren't.

1builder* not a team
3,389tests passing* CI, 2026-10-01, not a formal proof
6 → 0audits · criticals standing* self-run, not third-party; zero criticals as of 2026-09-25, core package
0known cross-tenant breaches in production* none reported; one leak found in audit and fixed, see below
1production instance* by choice, scale isn't the lesson yet

The system is honest because the person who designed it is. I made sure of that from the start, built in, not bolted on.

the invitation

The numbers hold. Asterisks included.

If you want a builder who attacks his own work this hard before anyone asks, and keeps doing it as the system grows, let's talk.