Expérimental~63 · IA en attente
My attempt at running a model: A short story.
r/LocalLLaMAu/fintip22 septembre 2026
Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.
Résumé
I was going to keep this brief, but then this came out. Felt like sharing the hardware explorations I've gone down. Some backstory, you can skip if you are more interested in hardware than I quit coding after many years right after chatgpt 3.5 became a thing. I tried writing a small shader with it for the first time--…
Afficher le post originalMasquer le post original
I was going to keep this brief, but then this came out. Felt like sharing the hardware explorations I've gone down.
Some backstory, you can skip if you are more interested in hardware than
I quit coding after many years right after chatgpt 3.5 became a thing. I tried writing a small shader with it for the first time--I'd done lots of work with 3d in javascript, but shaders (gpu code) were a dark art I had never touched. Spent a long weekend, back and forth, pasting in and out of the chat window.
In the end it was never able to accomplish it on its own. But through debugging its code and reading the docs and trial and error, I eventually learned enough through the process to actually get the thing running myself. A true, weird team effort where I got sucked deeper and deeper in until I was able to figure out why it was wrong and learned a lot along the way.
A very different experience than coding with AI now.
The problem was that I had been asking it to do the wrong things, and it was not able to tell me I didn't know what to ask for.
It got close on the first try, and every correction I'd tried to push it to was the wrong direction. I had to learn to correct myself.
I digress.
I quite coding for a couple years. Life stuff. Not relevant to this story.
I started doing some personal projects here and there this year, and started dipping my toes into AI, since that was apparently how it was done now.
I had a pro gemini subscription bundled for free with soemthing else, so started pasting code in and out of a chatbot.
It quickly became addictive.
Within a week, though, the limits of that context were causing real problems. it would drop parts on the output that I wouldn't notice until later, pieces of my website getting lost.
I realized I needed to use an editor.
Started using gemini plugin in vs code, my trusty tool for many years that I'd never quite loved as much as sublime code, but learned to live with.
That was another level of power. I started building a new project. It was incredible how powerful it was.
But eventually, cracks started to show.
I learned that it couldn't do architecture. It could consult on architectural discussions and give incredible insight, yes. It could also be a complete moron. I couldn't delegate tech choices to it, it would just pick the new hotness. It didn't share my values of cleanliness, simplicity. I went back and forth from treating it like a wishlist genie to not trusting it at all and checking everything and constantly finding issues with its choices and wondering what I missed when I wasn't looking.
Huge performance issues due to terribly written code.
Many days overhauling for algorithmic efficiency.
Struggles getting it to use Mithril, my preferred library over React, due to obscure edge cases in differences between the two that it would run into again and again.
Large scale refactors.
Days of ripping things out and rebuilding them. Repeatedly.
Shock at graphs just getting generated at one click. All those graphs I built by hand... knowing new devs would just... almost never do that again. I was the quaint old hand who did things the hard way in the old days, now.
...but then, constantly coaching it to not rebuild code from scratch, constantly demanding it re-use components, constantly seeing the same code show up again and again.
Eventually that plugin got deprected. I started trying antigravity, their own fork of vscode.
My first try was rough. It used up a 5 hour limit in about 1 hour or less of a test agentic task; I had never hit a limit in vscode with that extension. but the gui was slick.
a couple weeks later I tried again. incredible.
...for a few weeks.
...until I started going mad. The code had become more than gemini could handle. I was constantly having my hand in the code.
I wonder if I'll ever pay that much attention go code again, now?
Gemini wasn't getting better, and I could see everyone clearly saying claude was better, upset with the new models not coming out. This was 3.1 pro, which, they still haven't had another other than new flash models.
But also I Was getting sick of the limits of gemini, which I had started hitting there at the end regularly--they came much more often in the antigravity editor than they did with the vscode extension. it wasn't a scam, the results were better, but... at a cost.
I got a $20 plan for claude. Downloaded claude code. how does one really get work done in the terminal? but I was willing to find out.
...wow.
Instantly impressed. shocked. the interface was a joy.
but I had also started wondering more about local models. I had downloaded and toyed with the first deepseek local llm model I could run on my laptop--an rtx 4090m with 16gb vram. I got it to respond to an input on the command line. It felt magical and fascinating, but that was it, then.
I had even downloaded the very first version of stable diffusion years before and run that locally on the command line.
On reddit, I saw it evolving, things were moving fast. These context limits were a pain, and they kept getting tighter and tighter, while the local models sounded more and more capable...
Step 0:
I wondered... I had a 3080 laptop with 16gb vram and a 4090 laptop with 16gb vram... could I link them together and run something between the two of them? They also both had thunderbolt 4...
Well, after some amoutn of days of going mad scientist, I had experimental RPC mode running and 3.6 27b q4 running with solid context between the two, and shockingly, it produced usable numbers even in that configuration.
I kept using claude, but kept experimenting with this project in the background.
I started wondering: if I sold one or both laptops.... what could I build?
Step 1: sell a laptop. I decided to sell the 3080, after much dliberation. $1200, fb marketplace.
What to buy? I was sure it should be an rtx 3090, but then after considering a B70 longingly, I learned about the xtx 7900... and realized that it was a pretty solid 24gb option for a lot less money, and money was at a premium. And with thunderbolt 4 and an egpu dock... I could have 16+24 = 40gb of vram.
xtx 7900 24gb: $700. Could have had another at 600-650, but this one was closer.
TB4 egpu dock (AG02): $220.
That was an improvement, in speed and vram capacity.
Ran that for a while.
But kept wondering about a desktop... did that make more sense?
Step 2: The desktop.
I hadn't built a desktop since I was a kid 20 years before. I had only cursory awareness through random LTT youtube videos, mostly, about what it had become in the years since.
I had little idea what to buy, considered it on and off.
Finally one night, discussing it with claude for the 6th time or so, I finally felt like I had a conclusion.
Was it good? I still don't know, but I think so? For some reason I don't see many going this route, so you tell me.
X570 + 5800x included: $180 32gb ddr4 ram: $75 in-box-but-unused 1200w power supply (be quiet!): $80.
went to the nearby computer store and got a case for $75 that was the biggest cheap one they had, and a cpu cooler for $45 (there were cheaper ones, but this one was a bit quieter with a bigger fan).
I picked this specific model of X570 board because it can bifurcate the pcie lanes to x8 by x8 (aorus gigabyte pro wifi).
I plugged in the xtx. pwoer supply quirk meant I had 1 extra power supply pcie 6+2 plug available (3/4 to xtx, 1/4 left, realized I'd need a 12hvpr adapter to 3x pcie 6+2 ports for another card in the future). I found:
rtx 3060 with 12gb vram: $190
and decided that would work as an overflow card and a cheap temporary filler card while I sold the laptop.
I later got
64gb 3200mhz vram (16x4 kit): $275
but... I can only run it at 2990mhz, I learned through some trial and error. motherboard quirk. desktops are weird.
Still need to sell the 32gb and the laptop, and the fun part: buying a second xtx 7900 to get a 48gb machine that should run tensor parallel.
So, current desktop:
700 + 180 + 75 + 80 + 190 + 275 = 1500
• 220 for the egpu dock I'm not using right now, 1720?
When I sell the 32gb ram kit, let's call it 1600.
I did buy a used big wide screen for $120 too, I guess.
Remaining plan:
(1) sell the laptop. $2k. It's crazy how powerful a 4090 laptop chip is, it's a legitimate loss and nvidia has real benefits, but it's the tradeoff I have to make. Keep in mind the $1.2k I already got from the other laptop I had lying around. (2) sell the ram. (3) buy another XTX 7900. (4) nvme -> occulink adapter for the dock I have is in the mail. (5) get a thin, light, cheap thinkpad as a thin client.
I should just about come out cost neutral, depending on how much my 2nd laptop sells for and what I end up buying, and end up with a server with 48gb tensor vram, 60gb in a pinch when I'm willing to let the 3060 drag me down.
And if I want more? sell the 3060 and get a third xtx for 4-500 more.
A desktop with 72gb of decent mid-tier (?) gpu vram possible for what would be $2200~ all in?
3090s are nice, but I'd rather have two xtxs, they legitimately have been 1350 near me. I'll still take a cheap one if I find it, but that just hasn't been possible yet.
Does it make sense? I feel like it's the most efficient route to a decently high vram machine, but hopefully I didn't make a dumb mistake. Let me know if this was brilliant or dumb or somewhere in between, but my hunch is the price I spent for the quality tier I hit is worthwhile.
Sorry for being so verbose, perhaps the story was interesting to some other long winded soul.
