Automatisation~56 · IA en attente
I built an open-source tool to run 100% offline PDF data extraction and contract auditing using local LLMs (No cloud APIs, zero data leaks)
r/SideProjectu/PharwekCEO27 septembre 2026
Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.
Résumé
Hey everyone, If you’ve ever tried to pull structured data, clauses, or financial metrics out of messy 50+ page PDFs, you know the drill: • Uploading sensitive files to cloud APIs (OpenAI/Anthropic) is an instant compliance or NDA violation. • Writing regex or custom parsers for non-standard PDF layouts is absolute to…
Afficher le post originalMasquer le post original
Hey everyone,
If you’ve ever tried to pull structured data, clauses, or financial metrics out of messy 50+ page PDFs, you know the drill:
• Uploading sensitive files to cloud APIs (OpenAI/Anthropic) is an instant compliance or NDA violation.
• Writing regex or custom parsers for non-standard PDF layouts is absolute torture.
I got so tired of this bottleneck that I built a 100% local-first extraction pipeline that runs entirely on your machine using Ollama (Llama 3, Qwen 2.5, etc.). Zero data ever leaves your device.
What makes it actually work under the hood:
• Beating the "Lost in the Middle" problem: Long documents are automatically sliced into smart, overlapping chunks so the local model never misses critical details buried deep inside the text.
• Zero Hallucinations (1:1 Verification): Instead of letting the AI summarize freely, the pipeline forces it to output strict JSON schemas tied to exact verbatim quotes from the source file. If it can't quote the exact line, it doesn't count.
• Conflict Resolution: When overlapping chunks pull contradictory values (e.g., a fee update on page 2 vs page 45), the system flags and stores both instead of silently overwriting.
It’s slower than blasting data to a cloud cluster, but it gives you absolute privacy and 100% local control over sensitive files.
The entire project is open-source and free on GitHub: https://github.com/wessteq/revenue-auditor.
I built this to solve my own workflow pain, but I’d love to hear how you guys handle local PDF processing or what edge cases you've run into!
