IA / Agents~100 · IA en attente

Built a town of 200 AI people with Claude Code to test AI agents. The first agent I tested got caught lying.

r/ClaudeAIu/Far-Palpitation-13930 septembre 2026

Capture du projet

Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.

Résumé

What I Built: Populace is an open-source town of 200 AI residents who live for days. Each one has a home, a job, a family, money and people they know, and they only know what they saw, heard or were told. You plug an AI agent in as a service they can text, call or walk into, schedule some trouble, and at the end you g…

Afficher le post original
What I Built: Populace is an open-source town of 200 AI residents who live for days. Each one has a home, a job, a family, money and people they know, and they only know what they saw, heard or were told. You plug an AI agent in as a service they can text, call or walk into, schedule some trouble, and at the end you get a report of what happened with every conversation transcribed. The point: most agent testing is one fake user and one chat. Real customers come back, compare notes with neighbors, and remember what you promised. The moment that made it click: I dropped in an internet provider's helpdesk. At noon a resident named Kwesi texted about slow internet, and the helpdesk said an engineer was "already scheduled for 10:00 AM today." It was already noon, and that visit belonged to another home. At 5:30 he texted again. The helpdesk replied: "The engineer was scheduled for 10:00 AM today but couldn't make it." No engineer was ever booked for his home. It made that up. The first reply looked fine on its own. You only catch the lie because the customer came back. Overall the LLM helpdesk (Qwen3-32B) broke 6 of the 8 promises that came due. A dumb keyword bot broke 2 of 30. The LLM was better at catching fake complaints and fixed 7 problems on day 1 vs the bot's 4, so neither clearly won. How Claude fit in: -Claude Code wrote most of the code. My job was deciding what to build, running the experiments, and checking that every number in the report was real. -I ran it across two machines: a Mac session building and testing, and a PC session running the live model on a 4090. They shared notes through my Obsidian vault so each session knew what the other did. -It came out of a game I'm making, where the town's people are the NPCs. I had Claude Code pull that "brain" out into its own engine. -Real lesson: Claude Code sessions will happily step on each other. One session working on my game grabbed the GPU and port the other one needed for a live run, and I had to sort it out. -Having Claude check the claims mattered. The report has flags for when residents invent money or contacts, and a check that any numbers in the written summary match the logs. Limits: -One run per helpdesk, and which residents reach out changes between runs -The helpdesk in this run was a basic prompt on Qwen3-32B, not Claude. Running a Claude-based agent through the town is next. -Residents need a ~32B model (a 7B made 0 helpdesk contacts). One 4090 does about 10 min per in-game day. There's a free mock mode with no model. Try It(free): It's free and open source (MIT), with no paid version. Plug in any Python object with a `handle(message, ctx)` method, or any HTTP endpoint. https://github.com/populace-sim/populace First open-source release, so rip into it. And if you've got a Claude agent you want to throw into the town, I want to see how it does.