Finance~85 · IA en attente
Two days of load testing with Claude Code: p95 went 14s → 149ms. The first bottleneck was 42% of CPU spent rebuilding the same objects on every request.
r/ClaudeAIu/fyriyc29 septembre 2026
Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.
Résumé
I spent two days load-testing a multi-tenant POS platform (8 Node services on Kubernetes, MySQL-per-tenant, RabbitMQ, Redis) using Claude Code as the analysis loop. Went from a p95 of 14,250ms to 149ms on identical load, and ended with a measured capacity ceiling instead of a guess. Posting because the interesting par…
Afficher le post originalMasquer le post original
I spent two days load-testing a multi-tenant POS platform (8 Node services on Kubernetes, MySQL-per-tenant, RabbitMQ, Redis) using Claude Code as the analysis loop. Went from a p95 of 14,250ms to 149ms on identical load, and ended with a measured capacity ceiling instead of a guess.
Posting because the interesting part wasn't the speedup, it was the shape of the work, and the places the agent was flatly wrong.
The loop
Each cycle: I decide what to run → agent dispatches and waits out a ~50 min run → agent pulls k6 output, NRQL across 5 apps, kubectl, CloudWatch, V8 CPU flame graphs → forms a hypothesis → I review/correct/scope → agent implements across however many repos it touches.
We went round that 12 times. Twelve runs, twelve different binding constraints. Each had to be removed before the next became visible, which is why this kind of work can't be parallelised; you don't know what's behind the current wall until you take it down.
A few of the findings
1. 42% of CPU was framework overhead. First flame graph showed a middleware rebuilding the entire ORM model layer on every request; re-defining every model, re-wiring every association, producing a byte-identical result each time. Application code was 0.7% of CPU. GC was another 29%, mostly churn from that.
One WeakMap later: p95 14,250ms → 301ms, DB time 2,604ms → 50ms, 13 pod restarts → 0. It also dissolved two "data integrity bugs" that were just the async worker never getting scheduled.
2. Our APM agent was 38.5% of CPU. And this is the bit I'd want other people to know: a standard self-time-by-leaf flame graph view put it at 1.5%. The agent's work lands in frames attributed to node internals, so you only see it if you walk each sample's ancestry. Two config lines removed ~10% of CPU fleet-wide.
3. All five of our autoscalers were decoration. Every HPA used a 400m CPU target. Measured per-pod averages were 16–84 mCPU. They had literally never scaled on CPU.
4. Nodes were 95% "full" and 53% busy. The scheduler admits pods on requests, and ours over-reserved by 2.4–7.7× against measured usage. Meanwhile 15 of 27 pods had no CPU request at all — BestEffort, first to be evicted, invisible to capacity maths.
5. A five-minute error storm on every single scale-up. /health returned a static 200 and served as both liveness AND readiness, so it couldn't express "running but can't serve yet". The process started listening before secrets loaded. Every new pod took traffic and 500'd for ~5 minutes while its sibling served fine — so the service looked degraded, not broken.
The underlying reason was worse than a missing await:
static async getEnvironmentValues() { let interval = setInterval(async () => { ...fetch... }, 10000); // schedules, returns }
That resolves immediately. Four services did await it and were just winning a race.
6. One line of prefetch: 1*.* Switching from request-rate to invoice-rate load found a wall nothing HTTP-facing could see: every web metric looked great (p95 103ms, 0.18% errors) while the async confirmation worker fell behind 1.15/sec for 25 minutes straight and ended 7.2 minutes in arrears. Meaning stock-on-hand was overstated by ~1,700 invoices at peak. Invisible to every dashboard we had.
The failure that needed a human
At the top load the system collapsed with 12.85% errors, 10,930 create-failures. And everything the agent could see said nothing was saturated. App CPU at 34% of limit. Zero throttling. Query counts unchanged. But every statement uniformly slow, including a trivial indexed lookup going 5.7ms → 109ms, and Redis at 79ms, which has no locks and isn't even the database.
First hypothesis was lock contention. Fit the pattern, and was wrong.
The agent tried to pull connection counts. Its monitoring returned null. It said so, and explicitly labelled its own conclusion "strong but circumstantial."
Then said one sentence: "rds metrics can be in cloudwatch if you want to see."
That was the whole unlock. And the answer was the opposite of what someone expected:
connections DB CPU passing run 175/199 36.7% FAILING run 227/240 16.1% <- DB CPU FELL
The database was less busy during the failure. It was starved. 240 connections was almost exactly the app's own pool ceiling, against an instance allowing ~2000. The app had capped itself while the DB idled at 16%.
One config value: 25× latency improvement. 7,133ms → 283ms.
(The reason everything looked slow: the ORM measures a query from connection-acquire to result, so pool-queue wait gets billed to the statement. The queries never slowed. They were all standing in the same line.)
Where the agent was wrong
I kept a list. 14 corrections. The instructive ones:
• Claimed a config flag would save 3.8% of CPU. It already defaulted to that value. Measured change: zero.
• Concluded a fix wasn't deployed on two services; wrong container image layout assumption.
• Said a service needed more resources. It used 9 mCPU. Its problem was DNS lookups per request.
• Sized autoscaler targets from peak per-pod CPU when the HPA compares a rolling average. 7–12× too high. It reproduced the exact bug it was fixing.
• Divided one pod's CPU by the whole service's request rate. Internally consistent, directionally right, numerically wrong by 3×.
Two systemic patterns: it over-generalised from single observations ("this service has bug X" → "the fleet has bug X"), and it misread its own arithmetic. What made that survivable is that every claim came with the command that produced it, so a wrong conclusion was auditable rather than load-bearing.
Effort, honestly
Wall clock was dominated by the runs 12 × 50 min ≈ 10 hours of waiting no automation removes. What collapsed was everything around them. A senior perf engineer doing this manually: ~40–80 hours of skilled work over 2–3 weeks.
But the comparison isn't "40 hours vs 2 days." A human would have found fewer things, because rationing hypotheses is rational when each costs an hour. And they'd have known about CloudWatch on day one.
The actual shift: when testing a hypothesis gets cheap enough, you stop choosing between them. The human job moves from doing analysis to directing it and catching errors.
Where it ended
Proven sustained capacity ~14× our current production peak. And the thing that finally stops it is pleasingly boring: autoscaler replica ceilings pinned by a CPU-request budget on a small cluster, a third of which is consumed by workloads outside the system under test. Application CPU is 37%.
That's a capacity-planning decision, not an engineering problem. Best place a perf investigation can end.
Happy to go deeper on any of it, the flame graph ancestry-attribution thing, the readiness probe split, the pool-exhaustion diagnosis, or the "14 things the agent got wrong" list. I wrote most of this up properly; ask if you want it.
What's the most surprising bottleneck you've found that turned out to be config rather than code?