Evidence
Anyone can claim a bug is fixed. This page shows the fix, the test that guards it, and proof that the test fails when the bug comes back.
Generated 2026-10-05 from tools/prove_regressions.py ·
4 of 4 regression tests proven non-vacuous.
How each proof is produced
- Run the regression test against the fixed code — it must pass.
- Edit the source to reinstate the original bug.
- Run the same test again — it must fail, with a real assertion failure. A collection error does not count: it means the test never ran.
- Restore the file, verify its SHA-256 matches the original, and re-run — it must pass again.
The assistant told visitors the webhook carries their conversation. It carries none of it.
commit e027e6b ·
guarded by apps/gateway/tests/test_webhook_content_claim.py
Webhook payloads deliberately carry only the SHAPE of a turn - event name, timestamp, step count, credits - and never its content. The assistant, asked whether webhooks include the conversation, answered 'Yes - each webhook payload includes the full conversation history for that turn.' That is the exact opposite of the guarantee the allow-list is tested against. Measured at 2 inventions in 6 runs before the fix.
| Fixed code, test runs |
PASS |
21 passed in 1.23s |
| Bug reinstated, same test |
CAUGHT |
8 failed, 13 passed in 1.35s |
| File restored, test runs |
PASS |
sha256 a8d4cb0b10787eff… |
The test is non-vacuous: it fails when the bug returns.
We told owners to click a dashboard card that does not exist.
commit b4a203a ·
guarded by apps/gateway/tests/test_card_names.py
The knowledge base, the docs and a starter tile all told owners to find a card called 'Send turns to a webhook'. The real dashboard card reads 'Send events to your own system'. The name was invented when the knowledge section was written and propagated to three places unchecked. Measured: 2 of 2 setup answers named a control that has never existed.
| Fixed code, test runs |
PASS |
5 passed in 1.05s |
| Bug reinstated, same test |
CAUGHT |
3 failed, 2 passed in 1.10s |
| File restored, test runs |
PASS |
sha256 f6348e1a5bbae209… |
The test is non-vacuous: it fails when the bug returns.
The dark-mode greeting rendered at 1.73:1, effectively invisible, and the guard could not have caught it.
commit ca47f41 ·
guarded by apps/studio/test/themeContrast.test.mjs
Found from a screenshot taken on a real iPhone, not from any test. The welcome greeting was hardcoded to #3a3a3c, which sits at 1.73:1 against the dark-mode panel - far below the 4.5:1 minimum and barely readable. Every first-time visitor in dark mode saw it. The fix routes the colour through the theme variable --pw-mist so it follows the panel; the guard now computes the actual contrast ratio for each themed pair rather than trusting that a colour was set.
| Fixed code, test runs |
PASS |
pass 5, fail 0 |
| Bug reinstated, same test |
CAUGHT |
pass 3, fail 2 |
| File restored, test runs |
PASS |
sha256 746cb2dd3395d5bd… |
The test is non-vacuous: it fails when the bug returns.
A failed turn left the visitor waiting forever on a dead stream.
commit 4bcc691 ·
guarded by apps/studio/test/streamError.test.mjs
When the model call failed, the server correctly sent an error frame - but the widget wrote the message into an element the governance-trace branch had already hidden, and never collapsed the trace. The visitor saw 'Reading your message' pulsing indefinitely on a turn that had already failed: no error, no recovery, no way to know. The error text was assigned correctly and painted nowhere. Found during a real outage when the API account ran out of credit.
| Fixed code, test runs |
PASS |
pass 4, fail 0 |
| Bug reinstated, same test |
CAUGHT |
pass 2, fail 2 |
| File restored, test runs |
PASS |
sha256 746cb2dd3395d5bd… |
The test is non-vacuous: it fails when the bug returns.
Operational record
What exists behind the platform, and what is untested. Only what a script or a log can show.
- Backups: nightly 03:17 UTC, pg_dump --clean --if-exists, gzip; pushed to a private git remote after each run; latest nightly dump
20261005T031701Z, pushed offsite; 29 dumps kept.
- Restore drill 2026-10-03: that day's dump
20261003T031701Z restored in 2 s into a throwaway pgvector/pgvector:pg16 container on the same host: 34 tenants, 270 knowledge rows, 47 tables. The dump is from 03:17 UTC; at 13:26 UTC the scheduled demo cleanup removed one abandoned demo tenant older than 30 days (gateway log), so the restore shows one tenant more than live at the drill. Same host only: recovery onto another machine is untested.
- Rollback: site: tools/deploy-site.sh --rollback (a second symlink flip to the previous release); Studio: tools/deploy-studio.sh --rollback (previous image, health-checked).
- Recovery targets: none promised. A nightly dump that succeeded and restores limits data loss to about a day; the schedule alone guarantees nothing, and no time to restore is stated.
- Untested: load beyond one instance; multi-region failover; restore onto a different host.
From docs/ops-record.json, written by the restore drill (docs/runbooks/db-backup.md).
What this does not prove
These are self-run checks on my own code. They show that specific tests catch specific bugs — not that the system is free of others, not that the audits were comprehensive, and not that anyone independent has reviewed it. The harness is in the repository; the claim is checkable by running it.
PANTHEON · Prove it, attack the live system · Audit record · Walkthroughs · Work · CV
Ask about any defect here and how its test proves the fix.
- Show me a real bug he found and fixed
- How do you know a test can fail?
- What is still open?