MCP~56 · IA en attente
A ramble on my LumaBrowser project and the quirks I had to develop along the way to make this "Local Anthropic/ChatGPT in a box"
r/LocalLLaMAu/valdev17 septembre 2026
Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.
Résumé
I need an outlet to talk about my project and to brain dump the colossal amount of information I've collected along it's development. No one else understands it, and frankly I don't want to be that meme of the guy holding people captive while I talk at them about how creating a "singularity" mode for toggling between…
Afficher le post originalMasquer le post original
I need an outlet to talk about my project and to brain dump the colossal amount of information I've collected along it's development. No one else understands it, and frankly I don't want to be that meme of the guy holding people captive while I talk at them about how creating a "singularity" mode for toggling between LLM/Image/Editor/Video models on a single 5090 was a pain in the ass.
...and then how making improvements to that auto loading required making a RAM pinning solution for model files
...and how that got around using a RAMDISK because while Linux has that feature built in, windows really wants to have a driver to do that right
...and then while that got loading speeds to around 10 seconds, hot reloading was bottle-necked by the read speed of llama cpp so I made my own patched branch
...and then I didnt like how the context had to reload so I made improvements to be able to RAM pin the context for hot reload as well...
I think you get the point, and that's all just one small rabbit-hole.
What is LumaBrowser
A project that got out of hand. I've been building automation tools for decades now and I wanted to make a swiss army knife for doing handling things that went beyond needing Selenium or Playwright. So electron, was for once, the right choice for the automation platform.
My first feature was network interception, as I was running into an issue with a certain website not having the ability to create automated reports... but having the ability to show the report data in their reports area... capturable as a json request... So having a way to intercept those calls and forward them programmatically was useful.
Second feature was forwarding notifications, as I work with many companies and I am in too many slack channels, have too many emails, and am part of too many chats in general... And not all of them offer apps or APIs to capture chatter, but all of them do offer a web client that gives browser notifications... So, I built around that. (And now have a database that aggregates all of my notifications, and an AI client that reviews them, tags them, then sends me important messages grouped by context every 10 minutes, and otherwise groups the unimportant ones and notifies me once an hour). Shout out to ntfy, very useful for phone notifications.
Then I needed to add a selenium interface, playwright interface, all sorts of api backends for accessing and controlling tabs, as well as some algos for breaking down page content. Which led me into making the entire system have a base set of core functionality, and making everything else into a modular based approach/extension approach (Where the core is extended from, being the browser and other select features, but then other extensions can even extend off of eachother).
From here I got the bright idea of adding an MCP server that extended out this functionality, and adding overrides to the selenium/playwright/api connectors to be able to have LLM fallbacks for automation (thing "#submitButton", "the submit button") as a solution to html fragility. (fun note, it also reports back the corrected selector if the fallback is needed)
And thats... that was about the same time I was getting annoyed with LM Studio taking so long in getting in updates for llama cpp when new models came out. So I decided to build that in as well...
The AI side of things
A few years back I had made a fun little tool called "Talk to luma" which required me to train a small BERT model to contextually understand requests, then route requests to the appropriate functions or small AI's. The goal of this tool was to create a sort of dashboard that users could create on the fly, where all of its tooling came from the users request. It was meant to more or less be a demonstration of what my ResonantJs framework could do, and... was convoluted! So, I decided to try and implement llama cpp into LumaBrowser to get around needing to use LM Studio as its MCP system and get newer llama cpp builds in quicker... And... It worked perfectly fine, missing some bells and whistles, but it worked. My first tool here outside of the immediate self integrations was "AI Chat", which was/is a quick chat bar to interface with the current web browser tab, and it utilized reading the DOM, extracted minimal context, and other basic tools I made at the time. But I quickly realized, LLM's are ass at navigating web pages. Especially when the pages navigation is bad or the HTML is obscure (or long) or if the page is dynamically loaded. So I had to improve that with a mix of improvements to my navigation extractor and page content algo, but also... the template engine. Which is a self healing, reusable web templating engine that can be used across multiple lumabrowser instances. Which was great because without having a templating engine extracting only what the LLM needed to know to navigate a page, it was very context heavy. Around this same time I was struggling with loading larger models, specifically managing how it utilized multiple GPUs -- so I built set of tools to identify my machines setup (GPUs [with pcie version accounted for in ranking], RAM and their speeds), then analyze the models headers to determine distribution and how the model should load for the highest performance. This led into creating a fit tester for testing model + context length + context cache precision. I then added a few other supported interface runtimes and I was done.
Until I wanted image generation
I was using stable diffusion, then invoke, then comfy UI and had a bunch of scripts made to auto unload my LLM server when a request was made to my image generation server. This was annoying. So I played around with stable diffusion cpp and figured out it was easy to interface with, and my modularization approach and model identifier could be extended to make this work as an all in one. Fun side effect of this, it also allowed me to add extended tools to the AI hosted through lumabrowser to make images... And so... I made the LLM tab, which broke the AI chat away from the browser and now into its own area. This is where all of the fun begins with auto loading and unloading models comes in. I am fortunate enough to have a 5090 and a 3090 on my main machine as well as 192 GB of RAM, but even this doesn't always work right (or performantly) with the right combination of models. So I created an ability to cluster the models onto different cards (or span different cards in the case of the LLM)... Or to create a singularity group, which will auto unload a model for another model... if the other model is requested. (This is kind of my golden child, and its details are more written out in the beginning of this rant) At this point I started to just go wild, adding image editing, video generation support, music generation support but I realized something with all the tools Id created. It was making the system prompt stupidly long. Solution... make the system prompt dynamically populate only when the LLM needs access to a requested tool. And then I was done.
Until I wanted to edit code
And this one was a doozy, creating a whole seperate harness as an extension of the LLM chats existing feature set and adding guarding/approvals was nuts. Adding a full vs code editor in the mix was difficult, though easier is monaco editor. Then adding code artifact support, code validation support and everything else I had expected to do my job. (And since this I've also added a command line version for `luma` calls, as well as plugins for all jetbrains browser -- and I just got in a "check in notes" button as well, as I was getting sick of paying for github copilot for just this singular feature.) Then I was done...
Until I wanted to make 1000 other features... which Ill just bullet point out so I can save everyones time who decided to read my rant (which btw thank you)
• Ability to run multiple nodes on a network, connected to eachother to utilize eachothers gpus (works shockingly well)
• Added the ability to create a web server to access the lumabrowser llm remotely
• Added the ability to share conversations from lumabrowser remotely if the above is setup
• As well as sharing dynamic artifacts (like if you ask it to make a time tracker, you can share that code online, and yes... it can use its own db to keep sync)
• Added a dashboard area where you can pin those dynamic artifacts
• Added a full roleplay mode that can dynamically generate images for all characters, change background images and everything on the fly
• Voice support mode with multiple model sizes, this is for voice to text and text to voice.
• MCP server support, Lumabrowser already emits its own extensible MCP server but it needed its own ability to import them.
• Dynamic tool generation, the AI can make its own tools by researching API's or utilizing the web browser as a pseudo api layer itself.
• Agents, the goal of the above to be able to create focused AI agents with specific tools to get specific jobs done.
• Scheduled tasks, timed based tasks that do... well, whatever you want them to do.
• Ability to connect to a master lumabyte instance, utilizing its configured AI models. So if you have like a low powered laptop and want to connect to your server cluster but maintain autonomy of its own agents and what have you...
• Game mode, this one is new, the ability for you to create a game with the AI and have it auto generate image assets (if you have the image server configured), it runs/tests/verifies the game on its own and then lets you play it... OR share the link (if you have the system setup to have its web server exposed!). Oh, and then I upgraded it to also allow for "ai generated ai accessing" games, so like if you want NPCs to be able to respond using the LLM or want levels to dynamically generate using your LLM.
• Then tons of other things like NInfer support and auto setup on windows (utilizing wsl)
• Lots... and lots more...
Sorry for the unload for what its worth, but in the past... almost year there has been SO MUCH, and almost no one to talk to about it. This is literally the only community I could imagine who would either care or understand.