Homelab~45 · IA en attente
Kestra 2.0 is out. Easier to self-host and scale, still Open Source
r/selfhostedu/tchiotludo8 septembre 2026
Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.
Résumé
Disclosure: I am the CTO and co-founder of Kestra. It is Apache 2.0. Links at the bottom. We shipped 2.0 this morning. Most of the engine is new, and it is the version I always wanted to post here. Below is the road that got here, starting building alone during nights, then a first release, then four years of people t…
Afficher le post originalMasquer le post original
Disclosure: I am the CTO and co-founder of Kestra. It is Apache 2.0. Links at the bottom.
We shipped 2.0 this morning. Most of the engine is new, and it is the version I always wanted to post here. Below is the road that got here, starting building alone during nights, then a first release, then four years of people telling us what could be improved and us fixing it. Then what actually changed in 2.0, and what did not.
Where it comes from
In 2019 I was at Leroy Merlin, a big French hardware retailer. We ran Airflow and we lost tasks. Not tasks that failed, tasks that vanished because the scheduler itself fell over. A colleague told me "if you think Airflow is bad, do better". So I spent about thirty months of nights building a scheduler on Kafka and Elasticsearch. It went to production there in 2020, and in February 2022 I posted the first public release on /r/dataengineering.
The first thing people told me was that requiring Kafka and Elasticsearch to run a cron replacement was absurd. They were right. Four months later we shipped a JDBC backend, and one Postgres or MySQL was enough. That pattern, someone tells us the deployment story is bad and we fix the deployment story, is most of the history of the project. It is also how 2.0 happened.
https://preview.redd.it/gciuhogofboh1.png?width=2400&format=png&auto=webp&s=4de086b3ece72ae58b3be98a43f630d7fd9f981c
What it actually is
YAML flows. Each flow has triggers (cron, webhook, a file landing somewhere, an MQTT or Kafka message, another flow finishing) and tasks. A task is either a plugin (Postgres query, S3 copy, HTTP call, Ansible playbook, Terraform apply, ~2,000 of them) or a script in whatever language you want. Every script task runs in its own container by default, through the Docker socket you mount. If you do not want Docker on that box, the Process runner runs it straight on the host. On Kubernetes it becomes a pod. Same YAML in all three cases, one property changes.
You get a UI with logs, a Gantt view per execution, retries, replay from a failed task, and an editor in the browser with autocompletion from the real plugin schemas. Flows live in Git if you want, there is a sync task and a CLI for CI.
It is a JVM application. I will say that up front because I know this crowd. It wants 4 GB to be comfortable. Images are published for linux/amd64 and linux/arm64, and the arm64 build is the same one we ship for Graviton, so a Pi 4 or Pi 5 with 8 GB runs the server fine. Do not try it on a Pi Zero. If you have three cron jobs, keep cron.
Fastest way to look at it, embedded H2 database, nothing else needed:
docker run --pull=always --rm -it -p 8080:8080 --user=root \ --name kestra \ -v kestra_data:/app/storage \ -v kestra_db:/app/data \ -v /var/run/docker.sock:/var/run/docker.sock \ -v /tmp:/tmp \ -e KESTRA_PLUGINS_AUTO_INSTALL_ENABLED=true \ kestra/kestra:latest-slim server local
For anything you keep, use the Compose file in the docs with Postgres. Plugins download on demand with that env var, so you are not pulling a 2 GB image full of Snowflake connectors you will never use.
What changed since 2022, the parts that matter here
I will not list every release. If you looked at it years ago and left:
• Postgres or MySQL for everything. Queue, state, metadata. No message broker. Files go to local disk, or S3/MinIO if you want.
• Docker images compatible with ARM, tagged latest, latest-lts, or pinned v2.0.x. The LTS tag gets fixes for a year.
• Task runners. The same script task runs on Docker, on the host process, or as a Kubernetes pod. Switching is one YAML property.
• Realtime triggers. A flow can start within milliseconds of an MQTT message or a Kafka event, not on the next poll.
• Git sync, then a real CLI. kestractl ships since 1.3, with GitHub Actions for validating and deploying flows. There is also a Terraform provider if you would rather manage flows and namespaces as kestra_flow resources next to your Proxmox ones.
• Unit tests for flows, with fixtures, so a change does not have to hit your real NAS to be checked.
• Playground to run a flow task by task while you write it.
• Infrastructure plugins. 1.3 added ArgoCD, Cloudflare, Canonical MAAS, KVM, NetBox, Nutanix, Open Policy Agent. There is no Proxmox plugin. People do it with the Terraform task and the bpg provider, or hit the Proxmox API with the HTTP task. If someone wants to write one, I will review it.
• Kill switch, credentials, plugin versions pinned per flow. The boring stuff that makes it safe to leave running.
The first time I saw it used at home was a blog post last year: Proxmox VMs created by Terraform on a Git webhook, nightly rclone backups with retention, Ansible playbooks on a schedule, all in one place with logs. That made my week more than any enterprise deal did.
Why we rewrote the engine for 2.0
https://preview.redd.it/7e4ktjepfboh1.png?width=2880&format=png&auto=webp&s=cfe44a653c6a3ecbaaef5f45506632a58e2a841e
In 1.x every worker needed a connection to the central database. If you wanted a worker at the edge, say a Pi in the garage next to the sensors, or a box at your parents' house doing their backups, that Pi needed a route to your Postgres. Most people either exposed the database over a VPN or gave up and ran a full Kestra per site. We watched people give up more than once, at home and at work.
2.0 puts a controller between the workers and the backend. Workers are stateless, hold no database credentials, and open one outbound gRPC connection to the controller, with mutual TLS. Nothing connects in. So the control plane can sit on a small VPS and the worker can live behind a NAT with no port forwarding, no VPN, no database exposed. If the worker box gets stolen, there is nothing on it worth stealing.
That design did not come from a whiteboard. It came from a factory floor that allows no inbound traffic, a bank running jobs inside its cardholder network, and a hospital doing GPU inference next to its imaging archive. It happens that a homelab behind a consumer router has the exact same constraint, just with lower stakes.
The other debt: in 1.x, queue and repository came as a pair, and we kept two engine implementations to support it. Every bug fixed twice. 2.0 has one executor, one scheduler, one worker. You pick the queue and the repository separately. Start with Postgres for both, move the queue later without touching a flow. The Kafka Streams engine is gone and I am not sad about it.
What else is new in 2.0
• Loop replaces ForEach and ForEachItem. Each iteration is its own sub-execution, so a loop over ten thousand files can no longer eat the executor's memory and take the instance down. We saw that more than once.
• Trigger conditions are one when expression, same style as your CI pipelines, instead of a list of condition objects.
• Quotas cap executions per flow or namespace, so one broken schedule does not become everyone's incident.
• Large outputs load on demand. Thousands of tasks with gigabytes of outputs no longer slow down the UI or the database.
• Any flow can be an MCP tool. One trigger and the flow shows up as a named tool in Claude Code, Cursor, or any MCP client. The agent gets a normal execution, same logs, same RBAC, labelled system.from: mcp. An agent has no more rights than a person clicking Run, and destructive steps can wait for a human. If that is not your thing, do not add the trigger. Nothing else changes.
• The AI Agent task speaks to Ollama. Point it at the box with the GPU, prompts stay on your network. The Copilot in the UI can be turned off with one flag, kestra.ai.enabled: false, and then no AI endpoint exists on the server at all.
• New UI with a no-code canvas and Drafts. Canvas and YAML stay in sync. Drafts never run on a schedule until you publish. It is still always the YAML underneath.
What did not change
Declarative YAML. Any language, each script in its own container. The engine, the UI, the canvas editor and all the plugins are open source under Apache 2.0. We know major versions are where projects tend to change their license. We did not, and we do not plan to. The open source edition is what I run at home.
2.0 is an LTS: a year of backported fixes. The migration is mostly automatic, and there is a CLI that dry-runs the flow changes before touching anything. About forty teams ran the release candidates for the last two months, so the obvious breakage is already fixed. Tell us about the non-obvious kind.
Happy to answer anything about the architecture, the memory footprint, running it on ARM or k3s, or the parts we got wrong.
• GitHub: https://github.com/kestra-io/kestra
• Release: https://kestra.io/two-zero
• The homelab video I mentioned: https://www.youtube.com/watch?v=OePlSQGtbJs https://www.youtube.com/watch?v=PJG1-7hMHsE

