AI engineering · Field notes
The Agentic OS Bible
A practical guide to building an agentic operating system: start with one recurring job, give code the rules, agents bounded work, and humans the final call.
- Written
A lead comes in while you are on a job. They ask for a quote, mention a tight deadline, and attach a photo. Nobody replies until tomorrow. You do not need 20 chatbots with job titles to solve that. You need one system that catches the request, knows what it can do, prepares the next step, and tells you when something needs your decision.
That is what I mean by an agentic operating system. Not a new computer OS, and not a chatbot that promises to run your company. It is the set of rules, records, tools, workers, checks, and approvals that lets useful work continue after you close the chat window.
I have spun up 55 agents at once. It was technically impressive. It also left me spending time keeping them on task, so I made an agent to manage the agents. That experience changed the question I ask. I do not start with how many agents I can create. I start with which job should reliably get done.
Start with the job, not the roster
Take the missed lead. A useful first job is: collect new inquiries, identify the requested service and deadline, prepare a response, flag anything uncertain, and put a draft in front of the owner before it goes out. The system should record whether anyone approved and sent it.
Write down the trigger, the input source, the owner of the decision, the permitted output, and the proof of completion. If the inquiry never arrived, say no inquiry arrived. If the photo cannot be read, ask for clarification. If the draft is ready but nobody approved it, the job is waiting, not done.
This is where many agent demos cheat without meaning to. A polished draft is easy to show. A complete handoff requires evidence that the right lead was found, the details were extracted accurately, no duplicate reply was sent, and the human knows what decision is left.
Give each layer a job
The intake layer receives the email, form, or message. The record layer keeps the original inquiry, source, customer identity, status, and decision history. The tool layer reads that record and drafts or updates work through explicit interfaces. None of those parts needs to pretend to be a person.
Code owns the hard rules: who can see the lead, whether a message has already been sent, which provider may handle customer data, what costs money, and which action requires approval. A language model such as GPT or Claude handles the ambiguous part: what does the customer want, what is missing, and what should a helpful draft say?
A bounded decision model such as Jev can help with smaller choices in between: is this a quote request, a complaint, or a scheduling question? Does the message need urgent review? Is a cheap worker likely enough? It gets a list of possible answers; it does not write the email or own your policy. Its structured answer can still be wrong. If the decision matters, measure its errors and escalate uncertainty. It is optional, not the price of admission.[1][2]
The coordinator, or Chief of Staff, should still be a capable language model when the work is messy. It turns your goal into bounded work orders, notices when a result is incomplete, and explains the choice left to you. A worker is not made useful by its title. Give it a narrow role, the right tools, enough context, a spending boundary, and a test for success.
The final layer is you. You approve the consequential act: sending the offer, spending the budget, changing customer data, publishing, deleting, or deploying. You can delegate preparation. You should not delegate accountability by accident.
Keep the record outside the conversation
A chat transcript is useful context, but it is a poor customer ledger. Store the inquiry and its status in an ordinary database or an existing CRM. Preserve the original message, source link, timestamps, draft, approval, and final send receipt. Generate a readable view from that record instead of asking an agent to maintain the only copy of truth in a free-form memory file.
There is a useful community example in the Hermes archive: a contributor extracts structured signals, writes typed nodes and edges into SQLite, then generates a read-only Markdown vault for viewing. That is a design idea, not a turnkey Hermes feature or proof that read-only files create a security boundary. Hermes also has its own persistent memory and session search; neither replaces the business record that decides whether a lead was contacted.[3][4]
The same discipline applies when one agent hands work to another. The second worker does not inherit every conversation you had with the first. Give it the actual task, constraints, relevant files or records, and what a valid result looks like. A good handoff is an instruction someone else could execute without guessing which “it” you meant.[5]
For long documents, do not dump a hundred pages into the chat and hope. A community demo indexes a PDF by section and lets the agent choose which pages to inspect. The method may save context on that kind of task; the archive does not establish that it beats every other retrieval system.[6]
Make “done” mean something
For the lead workflow, “done” cannot mean “I wrote a nice response.” The check is whether the draft refers to the right inquiry, carries the actual deadline and service, avoids invented prices, and is waiting for the right person to approve it. After approval, the system must be able to show which message was sent and avoid sending it twice.
A completion contract makes the outcome, verification, constraints, and stop condition explicit. A machine-run test can enforce part of that contract, such as no duplicate send or a valid link to the source record. A reviewer can judge the draft against the evidence. Neither a passing test nor a second model saying “looks good” proves the customer was actually helped.[7][8]
Schedule recurring work with the whole assignment written down: what to inspect, the time window, what to ignore, where to deliver, what counts as complete, and when to stop. A reminder that says “do the usual thing” has no usual thing when it starts in a fresh session. The official Hermes daily-briefing guide calls this a self-contained prompt.[9][10]
Earn the next agent
Run the first version with one workflow and an owner in the loop. Record missed inquiries, draft errors, duplicate attempts, review time, actual provider charges, and whether the approved response went out. Compare that with the process you already have. Faster typing is not the same as faster service.
Then choose the next layer only if it earns its place. If deterministic rules handle the queue, do not buy a decision model. If a smaller worker produces acceptable drafts, do not pay for a frontier model on every message. If a case needs interpretation, spend the intelligence there. If a second agent creates more review work than it removes, you have not gained capacity.
The Hermes archive is useful for discovering patterns, not deciding what is proven. Its own methodology says Jev rankings judge supplied text, not verified outcomes; author-supplied impression counts are not public performance data, and a sourced story is not an independently tested workflow. The operating system you build should be held to a higher bar: can you show the work, reproduce the decision, and tell who was responsible when it mattered?[11]
My rule: start with the recurring job. Give code the hard boundaries, models the work they are good at, the record the facts, tests the narrow checks, and the owner the final call. Add another agent only when the first one can show you, without bluffing, what it did and what remains.