Autre~54 · IA en attente
claude vs antigravity vs 3.827b vs flash next vs bonsai vs swift... built my own quality bench test framework using my own git history, early results are surprising and shocking.
r/LocalLLaMAu/fintip22 septembre 2026
Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.
Résumé
A note: this entire thing is 100% human written, not even AI drafted or edited. So, enjoy. Or not. edit for a tldr that completed just after posting: ┌──────────────────────────────────────────────┬─────────┐ │ config │ T1-core │ ├──────────────────────────────────────────────┼─────────┤ │ Flash-Next Q2KXL (2-bit MoE)…
Afficher le post originalMasquer le post original
A note: this entire thing is 100% human written, not even AI drafted or edited. So, enjoy. Or not.
edit for a tldr that completed just after posting:
```
┌──────────────────────────────────────────────┬─────────┐ │ config │ T1-core │ ├──────────────────────────────────────────────┼─────────┤ │ Flash-Next Q2_K_XL (2-bit MoE) │ 11/12 │ ├──────────────────────────────────────────────┼─────────┤ │ Swift-27B Q4_K_L │ 10/12 │ ├──────────────────────────────────────────────┼─────────┤ │ agy / Gemini │ 10/12 │ ├──────────────────────────────────────────────┼─────────┤ │ Qwen3.8-27B Q3_K_XL │ 10/12 │ ├──────────────────────────────────────────────┼─────────┤ │ Bonsai-2-27B (ternary, 1.74 bpw, 5.6 GB) │ 9/12 │ ├──────────────────────────────────────────────┼─────────┤ │ Haiku 4.5, low effort, full scaffold │ 9/12 │ ├──────────────────────────────────────────────┼─────────┤ │ Claude Code default = Opus 5 + Fable advisor │ 8/12 │ ├──────────────────────────────────────────────┼─────────┤ │ Haiku 4.5, medium effort, full scaffold │ 8/12 │ ├──────────────────────────────────────────────┼─────────┤ │ Haiku 4.5, low effort, solo scaffold │ 6/12 │ ├──────────────────────────────────────────────┼─────────┤ │ Qwen2.5-Coder-7B │ 0/12 │ └──────────────────────────────────────────────┴─────────┘ ```
read ahead for more.
The big ticket benchmarks are cool and all, but for those of us running locally, there are a lot of tradeoffs to consider.
These things usually come in 3's, and this is a classic:
Fast? Usually goes with small. Also goes with less...
Smart? Usually big and slow, and the smarter, the less....
Context? Slows things down, and if you want more it's often at the cost of a dumber model.
So big model, small quant? or high fidelity smaller model?
The model with fast prefill but slow generation, or vice versa?
Some models with some knobs collapse rates at high context.
Not to mention:
Harness factors?
And how the harness interacts with context and intelligence and speed: how does your setup handle compaction?
And: how good is your chosen model and quant on your kind of work?
It's enough variables to drive you mad. There's way too much seat-of-the-pants rating going on, and I don't trust any of it.
Constantly hear poeple saying that anything below q4 is worthless, that q2 of 27b is actually great--no it's terrible--no it works good for me.
Is Bonsai worth anything?
Is that new swift model decent?
Does flash next not overthinking make it give higher quality for, in the end, the same wall time as 27b?
q4 k/v, or q8? or full-fat f16?
abliterated models--good, bad? better?
Does KLD tell the whole story?
...and every combination of the above?
...
So, I did want any sane dev who is losing his sanity could do.
I told claude to go through my git history--especially the last few months, but also all my code projects git history from the last decade that I keep on my hard drive for some reason, and look for bugs that would be good test candidates.
I was going to make my own test benchmark. One that nothing could have benchmaxed, because it's my own code. And projects that are the things I code (full stack stuff, javascript specialist, used to do 3d browser stuff, this last months lots of data science and app dev, etc.).
So I've been building that out. I have >100 candidates. I have 12 tight core cnadidates I've done the early work on, so full results are pending, because the first 12 testt take about 4-6 hours to run for local models on my machine.
Well, I will have a lot more to share in the future, if there's interest, but, the early results:
• Flash Next Q2 K XL has scored the best of any model I've tested, at 11/12. Most shockingly: even though it ran far slower, it was far more efficient with its tokens, so it actually did not take much longer than a Q3 of 27B--truly, very comparable times--while outperforming it.
• That Swift 27B that just came out genuinely did great, tied with two other models for second at 10/12*, ran fastest of my strong local models, and used far, far fewer tokens--about 1/3rd or so. This is likely going to be my daily driver.
• a q3 of 27b vanilla (unsloth) also was in the models that tied for 2nd with 10/12.
• Gemini 3.8 flash, default medium setting, scored 10/12. It uses its own agy-cli harness; all local models used opencode, btw.
• bonsai, which I haven't taken seriously? well, it scored a respectable 9/12
• Did you see the recent paper that I just saw a headline for a few hours ago about claude getting handicapped? Well... 24 hours before that, I was really trying to process the results I was seeing.
Claude code with open 5 + fable advisor scored 8/12.
Yeah.
And Haiku medium, likewise with fable advisor? 8/12 as well.
Haiku low, no fable? 6/12.
So... for what I can say, my early results seem to validate that.
I'll throw in the caveat that my test suite may be disproportionately made up of claude weak spots, as I have a lot of claude code from my most recent work, so take that with as much salt as you deem appropriate. hopefully the corpus of tests improves in time.
I tried a run with qwen coder 7b but it failed at tool calls so I need to get back to that one with an update I've had pending to get it working.
I won't be publishing my tests. they stay valuable because they're not public. But I Will be publishing my framework, with guides on how to prompt an AI to farm your own git history for your own tests. As you continue using it and find interesting test cases, throw it in to your personal bucket. You'll slowly build up a real repository of tests you care about, and as new models come along, you can test them on the metrics that matter to you.
Building this into the tuning and benchmarking tool I've had AI vibecoding since day 1 of getting qwen 3.6 27b running a few months back, so that'll be a thing too.
This isn't a promotion--if it was, I wouldn't be putting this post out there until I was ready to share it--but I do feel like this will close a really important part of the loop for local llm users like me in deciding what models to daily drive on metrics that matter in a much more substantial way.
Another bonus: it's perhaps a way to keep tabs on cloud models dropping quality?
Anyways, these preliminary results were too interesting to not share, so I figured I'd throw this out into the ether before I go to bed far too late at night once again.
Cheers.