Autre~100 · IA en attente

I benchmarked a 30-line Claude Code plugin on Opus 5.5 and Sonnet 5.5: half the cost per task at effort max, and more tests passed

r/ClaudeCodeu/Last_County67929 septembre 2026

Capture du projet

Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.

Résumé

Coding agents multiply everything: whole files read to change one line, logs pasted into context, 30 fixture files written by hand, the same check run three times. Every token in the context is paid again on every later turn. Occam is a Claude Code plugin that loads a 30-line rule set (about 650 tokens) at session sta…

Afficher le post original
Coding agents multiply everything: whole files read to change one line, logs pasted into context, 30 fixture files written by hand, the same check run three times. Every token in the context is paid again on every later turn. Occam is a Claude Code plugin that loads a 30-line rule set (about 650 tokens) at session start: build less, work lean, talk less. I benchmarked it on 9 coding tasks × 2 seeds, each solved without a plugin, with Occam and with Ponytail 4.10, in real headless Claude Code sessions. Hidden tests check every result. Cost per task vs. no plugin Model Effort Occam Ponytail 4.10 Tests passed, no plugin → Occam Opus 5.5 max −51% −26% 16 → 18 of 18 Sonnet 5.5 max −55% −9%* 16 → 17 of 18 Opus 5.5 medium −11% +11% 17 → 18 of 18 Sonnet 5.5 medium −3%* +26% 16 → 18 of 18 Haiku 4.5 – −2%* +30% 10 → 13 of 18 * within the noise: the 95% interval includes zero. All intervals are in the README. What I learned: • The saving grows with thinking. At effort max, thinking is almost half the bill. Occam roughly halves both thinking and turns (Sonnet 5.5: 16.5 → 8 turns per task). • At effort medium there is little to save, because the models hardly think. Occam still passed more tests there. • Rules cost context, too. Occam loads about 900 tokens per session, Ponytail about 3,100. On short sessions that alone costs Ponytail about 25%. • Favorite failure: asked to "clean up" four copy-pasted CSV exporters, both models without a plugin made the file longer (77 → 88 lines on Opus, 77 → 89 on Sonnet). Install /plugin marketplace add quisbaum-prog/occam /plugin install occam@occam /occam:occam lite|full|off switches the level. /occam:audit 7 shows where your own tokens went in the last 7 days. How I built it: with Claude Code. Claude wrote most of the rules, the hooks and the benchmark harness while I steered, and every benchmark round runs as headless Claude Code sessions. Limits: 18 tasks per round, mostly Python, single-prompt tasks. Costs are list-price estimates from the session metadata. Repo with raw data (metrics, final answer and code diff of every session) and the harness to rerun it: https://github.com/quisbaum-prog/occam Credit to Ponytail by Dietrich Gebert, which inspired the approach.