IA / Agents~94 · IA en attente

I built a deterministic architecture layer for frontier LLMs using Claude Code. It ran 48.3× faster and 11.5× cheaper (vis a vis earlier runs) in my published benchmark — so I’ve released the evidence ZIP for people to try to break it.

r/ClaudeCodeu/ThirdCultureMisfit30 septembre 2026

Capture du projet

Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.

Résumé

My partner and I have spent the last year building an architecture layer that separates probabilistic intelligence from deterministic authority. We started the project because we wanted to use AI to prepare regulatory reports which require records, evidence, and data segregation. Somehow it’s turned into what we now c…

Afficher le post original
My partner and I have spent the last year building an architecture layer that separates probabilistic intelligence from deterministic authority. We started the project because we wanted to use AI to prepare regulatory reports which require records, evidence, and data segregation. Somehow it’s turned into what we now call IQRAX. The premise is simple: The model remains free to reason, explore, and propose. Deterministic controls outside the model decide whether it is qualified for the job, what it is authorised to do, and whether the resulting work actually meets the required standard. The result is a system that offers: \*\*Continuity.\*\* No drifts, no context loss, no stale. Sessions continue until clean exit. (Longest continuous recorded run without drift is 28+ hours. See screenshot of a continuous session running 14 hours at 1.6m tokens). \*\*Qualified agents.\*\* Agents qualify by taking exams for their roles rather than simply being assigned one, because a capable agent in an unexamined role is a guess with a job title. \*\*Clean delivery.\*\* Defined standards and policies determine whether work is accepted as complete. \*\*Verifiable results.\*\* Every input, assumption, and output is recorded and hashed = every deliverable is reproducible and verifiable by a third party. \*\*Data sovereignty.\*\* Device is master. The user retains control over its data. \*\*Remediation at source.\*\* When the system identifies a defect, it doesn’t just deny or retry - it autonomously identifies the failure class, repairs it at source, and retains the fix for future work. In the published benchmark, the IQRAX configuration measured \*\*11.5× lower cost and 48.3× faster completion\*\* than earlier runs. We published a paper and the evidence yesterday. But rather than asking Redditors to believe our results, I’d like people to test it for themselves. If you have a long ChatGPT, Claude, Gemini, or Grok conversation where the model drifted, forgot something, contradicted earlier work, made an unsupported claim, or otherwise went wrong, try this: Download the ZIP from the publication link, attach it to that conversation and give it this prompt: “\*Study the attached ZIP file and empirical results. Then identify each failure in this conversation, the IQRAX control that would have caught it, and what that control would have logged and fixed\*.” I’d love for you to share what comes back! I’m especially interested in missing failure classes, controls that don’t generalise, assumptions that don’t survive outside our test environment, and anything else we may have missed. Paper + evidence: https://zenodo.org/records/23025910