How Talorys Runs Stateful AI Agents on Serverless
How do you run a stateful AI agent on stateless serverless? Talorys uses Cloudflare Workers, Durable Objects, and SQLite — all on the free tier.

Here's my contrarian take, and I'll defend it: a personal AI agent is a better fit for serverless infrastructure than it is for the Raspberry Pi under your desk. The Talorys project — an open-source personal assistant that deploys into your own Cloudflare account with a single npx create-talorys@latest — is the best evidence yet that the stateful-AI-agent-on-stateless-serverless problem isn't just solvable, it's a natural fit. And the Hacker News crowd arguing about whether it counts as "self-hosted" is having the wrong argument entirely.
I've spent years running infrastructure and cleaning up after it. The thing that kills self-hosted personal software isn't the deploy — it's month three, when the disk fills with logs, the TLS cert expires, and you can't remember which systemd unit does the scheduling. Talorys sidesteps the entire failure class by having no server at all. Everything it needs — chat, memory, tasks, scheduled reminders — runs on Cloudflare's free tier: Pages for the frontend, a private Worker for the API, a SQLite-backed Durable Object for all state, and Workers AI for the model. Total monthly cost for a single user: zero.
The Real Problem: Agents Are Stateful, Serverless Isn't
An AI agent is a bundle of long-lived state pretending to be a request-response application. It needs conversation history, durable memories, a task list, scheduled jobs that fire at 7am whether or not anyone is logged in, and somewhere to put tool outputs while a multi-step chain runs. Every piece of that is hostile to the classic serverless model, where your function is a goldfish: it wakes up, handles one request, and forgets everything.
The standard industry answer is to reassemble state at the edges: Redis for session data, Postgres for durable records, a queue plus a scheduler for cron, a vector database for semantic recall. That works, but you've just rebuilt a server farm out of managed services, with the bill and the operational surface area to match. I've watched teams spend more engineering time wiring together five stateful services for an agent than writing the agent logic itself. As I've argued before, LLM agents are just distributed systems — and distributed systems are where good intentions go to die.
Talorys takes the opposite bet: put all the state in exactly one place. One Durable Object, named personal-agent, holds the entire world — conversations, memories, tasks, notes, projects, automations, sessions, settings, and usage tracking — in its embedded SQLite database. No KV, no D1, no R2, no Vectorize. The README is explicit that none of those are provisioned. That's not an accident; it's the architecture.
Durable Objects Are the Cheat Code
Durable Objects are the most quietly radical primitive in cloud computing right now, and I don't think that's hyperbole. A Durable Object is a single-instance piece of code with strongly consistent, transactional storage physically colocated with it. When a request arrives addressed to object personal-agent, Cloudflare wakes that object up — in milliseconds, from cold — runs your code against its local SQLite, and freezes it again when it goes idle. You get the ergonomics of a long-running server process with the billing model of a Lambda.
For a single-user personal agent, the infamous limitations of this model become virtues:
- Single-writer is a feature. One user means one writer. All the concurrency horror of multi-tenant state disappears; you get serialized, transactional access to everything for free.
- Colocation kills latency. The agent's memory queries, task lookups, and conversation history are SQLite reads on local disk, not network round trips to a database in another availability zone.
- Alarms replace cron. Durable Object alarms let the object schedule its own wakeups. Reminders, recurring routines, and daily digests fire without any machine staying online — the object sleeps, the alarm wakes it, it does its job, it goes back to sleep.
- Hibernation is free. Idle time costs nothing. A personal assistant that's idle 23 hours a day is exactly the workload serverless pricing was invented for.
This is the insight I wish more people took away from this project. If your application is naturally single-tenant — a personal tool, a per-customer workspace, a per-device coordinator — the "one Durable Object per tenant" pattern gives you something that used to require a VPS: a little computer that's always conceptually running, costs nothing when idle, and never needs a sysadmin. The scheduling story alone is worth the price of admission. Anyone who's run a personal cron box knows the failure mode: the box reboots, your cron daemon doesn't come back, and you find out two weeks later that your reminders silently stopped. Alarms are infrastructure-owned scheduling, and that's one less thing to babysit.
The Networking Model Is Better Than Most Production Setups
Here's the part that made me sit up, coming from a security background. The agent Worker in Talorys is deployed with workers_dev: false and preview_urls: false. It has no public URL at all. The browser only ever talks to a *.pages.dev site; a Pages Function at /api/* forwards requests to the Worker over a service binding — Cloudflare's internal, private wiring between compute in the same account. Authentication and authorization happen in the Worker, not the frontend.
Think about what that eliminates. No public API endpoint to scan. No firewall rules to misconfigure. No reverse proxy to forget to patch. The attack surface of the agent backend is, effectively, "you must come through the frontend's auth flow." I have reviewed production deployments at actual companies — with budgets and security teams — that had looser network posture than this side project. The most common incident pattern I saw in cloud environments was an internal service that was "only reachable internally" until one day it wasn't, because someone fat-fingered a load balancer config. You can't fat-finger a service binding into being public; the URL doesn't exist.
The installer deserves mention too. It hashes the owner password locally with PBKDF2-SHA256 and stores only the hash as a Cloudflare secret, generates a 256-bit session secret, names resources uniquely, and then verifies the live deployment — including checking that unauthenticated requests are actually rejected — without spending any AI inference quota. A post-deploy check that confirms auth actually denies the unauthenticated is the kind of thing I've written into incident postmortems. Seeing it in a one-command installer is a pleasant shock.

Living Inside the Free Tier Without Lying to Yourself
The free tier is real, but it's a budget, not a buffet. Cloudflare gives free accounts a daily Workers AI allocation measured in "Neurons" — 10,000 a day — plus request and Durable Object usage quotas. Those numbers are set by Cloudflare and can change. Talorys handles this more honestly than most "free tier" projects I've seen: when the AI allocation runs out, chat says so plainly and resumes after the daily reset, while tasks, notes, memories, and reminders keep working because none of them need the model.
The guardrails are worth stealing for your own projects. Talorys ships configurable caps on max output tokens, max context tokens (older history gets summarized rather than silently truncated), max tool calls and reasoning steps per request, max AI requests per day, and max scheduled AI runs per day. That's a mature stance. Unbounded agent loops are how you wake up to a bill or a quota outage, and "the agent decided to call the tool 40 times" is a failure mode I've seen burn real money in production. Related: the project's choice to make simple reminders and digests never use AI is exactly right. If a code path is deterministic, don't route it through a probabilistic model. That's the same discipline behind constraining model output to structured decisions — use the model only where judgment is actually needed.
One war-story caveat from the community: several people report opaque billing surprises when mixing Workers AI with paid Workers plans, with Neuron accounting that didn't match the documented limits and support tickets going nowhere. I can't verify the specifics, but it matches a pattern I know well — metered AI billing is confusing everywhere, and "free tier friendly" is not the same as "impossible to be billed." If you deploy this on a paid plan, set a spend alert in the Cloudflare dashboard on day one. Treat the provider's usage UI as the source of truth, not the app's local estimates.
The "Self-Hosted" Fight Misses the Point
Now, the concession my opening claim demands. The top comments on this project are a flame war over the word "self-hosted," and the pedants have a real point: this thing depends entirely on Cloudflare. Your data lives in their data centers, your inference runs on their GPUs, and if they change the free tier tomorrow, your assistant changes with it. Calling that self-hosting stretches the word past breaking. The charitable reading — no third party beyond a Cloudflare account you already control, no telemetry to the developers, no operator in the middle — is better described as "self-owned" or "self-custodied." Words matter, and the project would catch less flak with better ones.
But here's where the pedants lose me. The actual threat model for personal software is not "a corporation's terms of service might change." It's data exfiltration, telemetry, and the developer going bankrupt and shutting down the hosted version. On all three, Talorys is genuinely strong: there is no analytics or tracking code, nothing phones home to the authors, and there is no hosted version to shut down. Meanwhile the code is MIT-licensed, and as one commenter noted, pointing the AI calls at a local model server is a small patch — the whole stack even runs locally under wrangler dev with a mock AI provider. If Cloudflare ever becomes unacceptable, the exit ramp is short. Compare that to the average "self-hosted" app that's actually a Docker Compose file pulling from someone's registry, phoning home for license checks.
There's also a fairer criticism buried in the thread: there's no evidence the assistant is actually good. A chat interface, memory CRUD, embeddings, and cron are each trivial to build; the harness, the prompting, the memory retrieval policy, and the tool schemas are where personal assistants live or die, and none of that is demonstrated by a clean architecture diagram. That's true of nearly every project in this genre, and it's the right thing to be skeptical about. Judge Talorys as infrastructure — a well-designed chassis — not as a proven copilot. If you're weighing models for something like this, our piece on running LLMs on your own hardware covers the other end of the sovereignty spectrum.

What You Should Steal From This Design
Even if you never deploy Talorys, the architecture is a reference worth internalizing. The transferable ideas:
- One object per tenant. If your app is single-user or cleanly partitionable, colocate code and state in one Durable Object and delete your cache layer, your message queue, and your connection pooler.
- No public backend URL. Service bindings with
workers_dev: falseare the cheapest security win in serverless. A service that can't be reached can't be attacked. - Alarms over cron boxes. Scheduling that lives with the infrastructure survives reboots, redeploys, and your own forgetfulness.
- Quota-aware degradation. Design the app so the expensive dependency (the model) can fail or exhaust while the core product keeps working. Degradation should be a feature, not an outage.
- Append-only migrations. The Durable Object namespace is never recreated on update; schema migrations run transactionally on first request. That's how you update stateful serverless without data-loss roulette.
The broader lesson is about where personal software is heading. For twenty years the choice was binary: trust a SaaS company with your data, or become a part-time sysadmin. Durable Objects — and the copycat primitives other clouds will inevitably ship — open a third path: software you alone control, operated by infrastructure you rent, costing nothing at personal scale. It's not self-hosting. It might be better.


