Autre~94 · IA en attente

Introducing Curia

r/ClaudeAIu/bobo-the-merciful29 septembre 2026

Capture du projet

Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.

Résumé

TLDR: Wrote a system for kicking off Claude Code sessions that have clear roles, responsibilities and memories. Laws and rules emerge over time. When paired with beads and a big backlog, this system allows for long running autonomous work across multiple Claude Code accounts without sacrificing quality. Automatic hand…

Afficher le post original
TLDR: Wrote a system for kicking off Claude Code sessions that have clear roles, responsibilities and memories. Laws and rules emerge over time. When paired with beads and a big backlog, this system allows for long running autonomous work across multiple Claude Code accounts without sacrificing quality. Automatic handoff between accounts and management of Claude Code limits is built in. The longer detail below is written by Claude and reviewed by me. Hi all, I've spent the last month building and using something I'd like to show you. It's called Curia, and it runs a team of Claude Code agents as a small society instead of a pile of independent terminals. Every time you spin up Claude Code, a named "seat" with its own memory and job is activated. Only one seat can be activated at a time, but multiple seats can run on one Claude account. Offices keep order, and written rules are enforced in every session. It's open source, and the four-minute video attached explains it better than I can in text. Why I built it By August I was running four to ten Claude Code terminals at once on a client engagement, jumping from one to the next all day. Solidly in stage 7 of Steve Yegge's 8 stages: \"The 8 Stages of Dev Evolution To AI\" - Welcome to Gas Town by Steve Yegge, January 2026 https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16dd04 (https://preview.redd.it/ygmc0nikagsh1.png?width=1412&format=png&auto=webp&s=d49e900dedca008b52f09b59ab7098e9fdadabba) I had already laid some groundwork by this point: http://CLAUDE.md files, skills, and Beads for the backlog. It worked, but I'd hit a ceiling on throughput. I was the bottleneck, and the same problems kept coming back: • Shared knowledge isn't memory. http://CLAUDE.md, skills and Beads gave every session the conventions and the backlog. Beads has a memory function for broader context - like Claude's in-built memory. But no agent remembered what it (i.e. it as a particular named agent) had been doing or why. Every session that wakes up was independent and had semi-amnesia. • Accounts ran dry at 2am. This was a problem I occasionally ran into - I was always juggling my usage limits across accounts and worrying about whether I had given an agent too much work to do over a nightly run for example. • Agents sat idle. They spent an hour watching a CI run they could have handed off, and every message between them went through me. • Rules were requests, not checks. Some lived only in my head, and one written in http://CLAUDE.md had nothing to stop a slip. Then Steve Yegge published a run of essays in August about running around fifty agents building his game Wyvern - his system is called Wheelhouse. His argument is that agent harnesses either devolve into chaos or evolve into societies, and what makes a society is law. His system grew seats, handoff notes, offices and "fences" that refuse what the rules forbid. It's closed source and built for his game, and he argues nobody else's could be reused, because every city's law is its own. Curia is 100% inspired by Steve's work on Gas Town and Wheelhouse. It splits the reusable machinery (one Python file) from the law, which stays in each project's own folder and gets written and built on for that project specifically. https://preview.redd.it/bejn2pinagsh1.png?width=1318&format=png&auto=webp&s=6f9eff13d4772c5a7f241ee740181f3c33a8e9f1 Curia itself is the same for everyone; each project's roster, rules and memory live in that project's own estate. Seats, offices, crew and fleet Every agent in Curia holds a seat: a named role with its own memory that outlives any single session. An office is one kind of seat. There are three kinds, and the difference is who wakes them and why: • Offices have a standing job and are woken on a schedule or when there's work. My Consul runs delivery, Notarius turns meeting notes into tickets, the Censor fixes a failing branch and the Lictor reports on the backlogs. • Crew are the seats I talk to directly, each with its own area, such as design, review or documents. • Fleet seats are workers an office sends out, each in its own copy of the repository. I never launch them myself. https://preview.redd.it/pxg125dqagsh1.png?width=1352&format=png&auto=webp&s=431b205c4955d63e1bbd1d0543f474375922a6ca I give the Consul its orders; it asks the crew for design and review and sends the fleet to build. A session is one waking of a seat, like a working day. The seat is what carries on: https://preview.redd.it/mmj82d1sagsh1.png?width=1318&format=png&auto=webp&s=9a23ba52b0d9a453f00b8126fea418c48361b917 Each session reads the note the last one wrote, so the seat picks up where it left off, even on a different account or model. What it has done so far On my main client engagement, which has five repositories, Curia nearly doubled my best by-hand pace, and nine in ten sessions ran with nobody watching. It took over on 30 August. These compare the four weeks before with the 24 days after: https://preview.redd.it/85hxkx2wagsh1.png?width=1368&format=png&auto=webp&s=89c0e6b24bd51bd09cd23c1f95244e4afec37491 Tickets are closed backlog items, not counting duplicates and tidy-ups; pull requests are counted from git. For context, back in the spring I was delivering about 25 tickets a week. Most of the climb to 149 happened in August, by hand, before Curia existed. What Curia changed is that the pace roughly doubled again while my job shifted from babysitting terminals to gathering context during the day, setting direction in the evening and reading reports in the morning. The mechanism itself is a month old - I've tweaked and built on it since: one Python file of about 7,700 lines, 162 tests, and no Python packages to install. What it can't do (yet) • It's Claude Code only. I have a design for other harnesses but not built yet. • It's young and opinionated. It's a month old and mostly proven on one client estate, by one person. Expect it to change. • It assumes my toolchain. Git, GitHub (the gh CLI), Beads for the backlog, Python 3.11+, and a Mac: the scheduled jobs use launchd. • It isn't a sandbox. Seats run on your machine with the tools you allow them. Fences are polite refusals backed by written rules, not a security boundary. • It's token-hungry. Over those 24 days, Claude Code's own cost reports add up to about $10,400 at list prices for me. • Night output depends on the backlog. It needs well-specified tickets. In mid-September it simply ran out of work, and anything waiting on a person's decision is left for the morning. • People still own the hard calls. Releases, client or critical decisions and anything labelled for a named person stay with a human. Review is done by other agents, not by me. Bug tickets filed have more than tripled, and most appear to be raised in review before merge. How I use it day to day https://preview.redd.it/0izjhdhyagsh1.png?width=1312&format=png&auto=webp&s=a4acfbab16880eb8457fa74d8b759fd5b268c60b During the day I talk to clients and the work gets filed; overnight the Consul builds it; in the morning I read what happened and decide what ships. In detail: • Talk to the client. I record our calls about features, and the notes land in my notes vault. • Notarius files the work. curia ingest <notes> wakes Notarius, the intake office, which turns the call into tickets in the right repository. Anything that needs a person's answer is labelled for that person, so the night leaves it alone. • Daytime is for thinking. I talk to the crew seats for design, review and documents when something shouldn't wait. • Evening: the Consul takes over. I launch it with curia launch consul --loop and tell it what to land or I just specify --everything. I can also set an end time - handy if I just want to run it for a few hours while I head out - this can be done with curia launch consul --everything --until 16:00. Overnight it has a design seat write up anything that needs it, then dispatches a fleet of six workers, each in its own git worktree. Another seat reviews each change, and the merge happens once it's reviewed and passing. A night can also be capped by a deadline or a budget. • 6am: the checks run. The Censor reads every integration branch and wakes a model only where one is failing. The Lictor reads every backlog hourly and nudges, but changes nothing. • Morning: I catch up. I read the handoff notes, answer the questions labelled for me, and decide what gets released. If any of this sounds like your week, have a look. The README covers installing it and starting with a single seat, and I'd love to hear what works and what breaks: GITHUB LINK: http://github.com/harrymunro/curia Landing page: http://curia.build Cheers, Harry