Routines: verified recurring work for small teams
We compared Hive against thirty agent products in September 2026. Nobody pays for orchestration anymore. They pay for one recurring job that runs correctly, gets checked, and never blows the bill. That is the feature we're building.
We spent 18 September 2026 reading the competition: canvas tools, agent frameworks, “AI
employee” SaaS, agent control planes, the closed vendors, and the local chat apps. Around
thirty products. We also re-read Hive’s own code — flows.py, rooms.py, llm.py,
triggers.py — to see honestly what we had and what we didn’t.
The conclusion fits in one paragraph. The 2026 market does not pay for orchestration. The open-source core is free everywhere — n8n, Dify, CrewAI, LangGraph, Paperclip. What people pay for is a repeated job that runs correctly, with someone checking the output, without wrecking the invoice. Hive already had most of the parts. The work is not another subsystem; it is wiring the existing pieces into one thing a small team would actually buy. We’re calling it a Routine.
What the market is complaining about
Read enough G2, Reddit and Hacker News threads and the same four complaints keep surfacing.
Cost surprises. Zapier sits at 1.4 stars on Trustpilot, largely over billing (see the Zapier Agents complaints). Lindy’s credits are opaque to users. Make charges 50 credits per agent run. Gartner’s token-usage report describes the same surge from the buyer’s side.
Brittle agents. CrewAI users ask how to stop infinite loops eating tokens. n8n’s AI Agent node has a memory failure thread long enough to be its own documentation. LangGraph shipped a self-hosted flaw chain in June.
Nobody checks the output. This is the quiet one. Almost every product runs the agent and hands you the result. Human-in-the-loop, where it exists, is a gate on a tool call (n8n 2.6’s HITL tools), not a second model reading the deliverable and saying whether the numbers are right.
Self-hosting is hard. Dify and LibreChat mean a stack of containers (a six-month Dify review is blunt about it). Open WebUI changed its license. OpenClaw exposed 21k instances to the internet.
Two more signals told us where the floor and ceiling are. OpenAI’s AgentKit is shutting down on 30 November 2026: a closed orchestrator from the biggest vendor in the space did not survive a year. And Paperclip went to 30k GitHub stars in three weeks on one idea: agents organised like an org chart, each with a budget that pauses the agent at 100 %. People are not starving for another framework. They want to know the bill will stop.
On price, Claude Cowork, Claude Code routines and Managed Agents set the frame: a routine is a recurring job with a deliverable, and 20–30 USD per seat is roughly what people will tolerate. Above that they self-host.
What the research says about “more agents”
There is a second thread, from the papers rather than the forums. The Cost of Consensus (arXiv, May 2026) and the earlier problem-drift work both find that free, homogeneous debate between models performs worse than isolated self-correction — the agents drift off the problem together. The pattern that wins is a verifier triggered on condition: one producer, one checker, and a heavier process only when the checker objects.
That matters to us because Hive has both shapes. Rooms run a /goal loop (producer, then
checker, two rounds) and a full council (four phases, 2 + 2N model calls for N seats, up to
six seats). The research says: make the first one the default, and stop treating the second
one as the grown-up option.
What Hive already had, and the gaps
Being honest with the code, the pieces were there:
- A checker in
/goal, and two parallel reviewers with mediation in/review. - Two-layer budgets: per agent (5 USD), per workspace (20 USD/day), plus per room, per project and per API key, with a reserve/settle ledger.
- Triggers: cron, HMAC webhooks on
/api/hook/<token>, email and ICS pollers, room schedules on a 30-second tick, OAuth apps for GitHub, Linear, Notion, Sentry and Cloudflare. - Deliverables: docx/pptx/xlsx generation, SHA-256 signoff manifests, Telegram out.
- Bring-your-own endpoint: any OpenAI-compatible
base_url, so Ollama, LM Studio or vLLM already work; embeddings already run locally onnomic-embed-text. - Evals to lock behaviour.
And the gaps, which turned out to be plumbing rather than architecture:
- Triggers land in Chat or Origin, not in a room’s
/goal— so a scheduled job never gets a checker. - The checker returns prose, not a structured verdict, so nothing downstream can branch on it.
- Council always runs the full three rounds, even for an easy question, and both council and
delegate sub-calls are marked
_no_cache, so the two most expensive paths never hit the response cache. (We expected to find missing Anthropiccache_controltoo; it turned out our gateway, cli-proxy-api, already injects the breakpoints for Claude, so that one was a false alarm.) - No outbound webhook when a run finishes. No Telegram in. No local preset or health check.
- Nothing packaged as “one job” you could point a customer at.
The Routine
A Routine is one recurring job, end to end, under a budget that stops itself:
trigger → room
/goal→ producer on a cheap or local model → checker on a strong model → conditional escalation → deliverable with signoff → delivered to the team’s channel
The customer-shaped version: “every morning at 8, collect Linear tickets and GitHub PRs, write a one-page brief, have the checker verify the numbers, post it to the team Telegram.” Or: “order form webhook → extract fields → match against the price list → export xlsx → wait for accounting to approve.”
Step by step:
- Trigger. Cron, HMAC webhook, email/ICS poller, and later Telegram in. The payload is
attached to the run as an input with a SHA-256, which
attach_inputalready does. - Produce on the cheap tier. The producer runs on the lowest rung of the model ladder,
preferring a local endpoint.
route_down, today an opt-in that only applies to the first attempt, becomes the default for Routines. - Verify on condition. The checker runs on a strong model in
taskprompt mode and returns a JSON verdict:pass | revise | escalate, a list of faults with quotations, and a confidence. Onrevise, the producer receives only the verdict and the rejected passage, not the whole draft — roughly half the tokens of the current second round. Two rounds maximum, which is alreadyGOAL_ROUNDS. Onescalate, or two rejections, two arbiters who were neither producer nor checker rule on exactly the disputed criteria — blind, in parallel, one call each. A defect clears only when both mark it ok with evidence; the merge is deterministic, so there is no extra Lead call. No cross-examination. - Hand over. The deliverable is a file in the room, signed off against a manifest, with docgen if the output is a document. The human gate is the existing 300-second tool approver on Today, later a Telegram approve button.
- Deliver. Telegram out (exists), plus an outbound webhook —
POST url {run, status, files, cost}— so the customer’s n8n, Make, Zapier or Google Sheet receives it, and the OpenAI-compatible/v1API so their existing tools can call Hive. - Budget that pauses. Each Routine carries a monthly
budget_usd. Hitting the ceiling pauses the Routine, notifies Telegram, and makes no model call. This reuses the ledger that already reserves and settles spend.
The cost comparison
This is the whole argument in numbers.
A full council with six seats costs 2 + 2N = 14 model calls on a strong model, every
time, regardless of how hard the question was. A Routine’s expected path is 2–3 calls: one
producer call on a cheap or local model, one checker call on a strong model, and occasionally
one revision. The worst case — two revisions and an escalation to the two-seat mini-council —
is about 7 calls. The heavy lifting happens on the cheap tier; the strong model reads, it
doesn’t write.
Around that, the token work, in order of payoff per effort:
| # | Change | Effect |
|---|---|---|
| 1 | Keep the stable prefix (harness, skills, catalog) ahead of the volatile board so gateway prompt caching keeps hitting | 60–90 % of system input tokens on every room turn |
| 2 | Drop _no_cache on delegate sub-calls (the key already covers the specialist’s own prompt) | Cache hits on repeated tasks |
| 3 | Structured verdict + a fix round that lists only the failed criteria | About half the tokens on /goal’s second round |
| 4 | Parallel delegate fan-out (a thread pool, like the room “wave”) | N× faster delegate workflows |
| 5 | A local rung on the ladder; route_down covers the /goal producer’s first draft | Producer close to 0 USD |
| 6 | Council rounds=1: blind positions → verdict, no cross-examination | Council down to 2 + N calls |
The rule we’ll document: local for production, extraction and classification; cloud for judgment and verification. Hive’s Brain (smart engine, graph, Origin) still needs Claude and stays that way. Routines don’t go through the Brain.
What we are not building
The market told us this as clearly as it told us what to build.
No drag-and-drop canvas — n8n, Dify and Flowise have won that, and Hive composes work with Rooms, not nodes. No fourth orchestrator next to CrewAI, LangGraph and Microsoft’s Agent Framework (AutoGen is already in maintenance mode). No seventy connectors. No agent message bus or presence. No council as the default. No local Brain. No marketplace, seat billing or multi-tenancy. A2A — 150 organisations and counting — waits until a customer asks for it. Wrapping Hive as an MCP server, so Claude Code, Cursor or n8n can call a Routine as a tool, comes first.
Sequence
Three waves, each locked by an eval before the next starts.
- Local presets (Ollama, LM Studio, vLLM) with a health check; a local rung on the ladder;
triggers that target
room:/goal; outbound webhook. Acceptance: a real cron runs/goalon Ollama, the checker on cloud, a file is produced, the webhook receives the payload. - Structured verdict and a fix round scoped to the failed criteria; arbitration on dispute; parallel delegates; cache on delegate sub-calls. Acceptance: token counts before and after on five eval cases; cache hits > 0 on delegate; a two-seat council in ≤ 5 calls.
- Telegram in with an approve button; per-Routine budget pause and notification; a Usage page per Routine; a landing page that sells Routines rather than seats. Acceptance: hit the ceiling → paused, zero model calls; a restore drill passes.
Why this fits Hive
Hive is a single process of standard-library Python and SQLite, self-hosted on your own box, with a Preact frontend that needs no build step. That is the answer to the fourth complaint before we write a line of Routine code: one process to back up, one file to restore, no container stack. The other three — cost, brittleness, verification — are what a Routine is for: a bounded job, a cheap producer, a strong checker that only escalates when it disagrees, and a budget that pauses itself instead of surprising you.
The market is paying for operational convenience and governance, not orchestration IP. That is where we’re pointing Hive.
Want something like this built for your team?
Get a quote →