Autre~80 · IA en attente
I built a Claude Code skill that makes Claude prove its bug fixes: every test it writes has to fail without the fix
r/ClaudeAIu/Sanechka_SS29 septembre 2026
Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.
Résumé
Claude Code fixes a bug, adds a test, CI goes green, and it says "done". But a green test only says the test passes. It doesn't say the test would have failed before the fix. So I built Receipts. It runs every test a change adds or edits twice: once with the change, and once with the changed source files reverted to m…
Afficher le post originalMasquer le post original
Claude Code fixes a bug, adds a test, CI goes green, and it says "done". But a green test only says the test passes. It doesn't say the test would have failed before the fix.
So I built Receipts. It runs every test a change adds or edits twice: once with the change, and once with the changed source files reverted to main. A test for a fix has to fail without it.
• PROVEN: fails without the fix, passes with it. That's the one you want.
• THEATER: passes both ways. It would have passed before the fix too, so it proves nothing.
• WEAK: fails without the fix only because it imports something the fix added.
It ships as a Claude Code plugin with a prove-fix skill. After Claude fixes a bug, it runs the check on its own tests before saying it's done. In a real session, Claude wrote a test for is_leap(2020), got THEATER (2020 was never broken; 1900 was), wrote a test for 1900 instead, and got PROVEN, without touching the fix.
/plugin marketplace add syntaxixr/receipts /plugin install receipts-check@receipts
To see how often this matters, I ran it over 100 pull requests with coding-agent fingerprints (mostly Claude Code) in Claude Agent SDK, OpenAI Agents SDK, the MCP Python SDK, fastmcp and simonw/llm, plus 81 maintainer fix commits in libraries like click and sqlparse:
• 82% of agent PRs and 90% of maintainer fixes were proven. Most tests do their job.
• In 10% of agent PRs, every test failed on the old code only because the test file imported a name the PR added at the top. On the old code the file can't even load, so nothing ever ran against the old behavior. No maintainer commit did that.
Example: claude-agent-sdk-python #1016 (https://github.com/anthropics/claude-agent-sdk-python/pull/1016). Its test file imports TaskUpdatedMessage and TERMINAL_TASK_STATUSES at the top. On the old code the whole file fails to import, so all 10 tests "fail", including an existing test for unknown subtypes. The fix may well be right; the tests just can't show it.
There's also an opt-in Stop hook that won't let Claude end a turn while its changed tests prove nothing. It's off by default, because it runs the repo's tests.
No LLM in the check itself: it's just your pytest, vitest or jest, run twice. Repo, GIFs and the full study: https://github.com/syntaxixr/receipts
Happy to hear where the verdicts are wrong for your code.