← All posts
HiveAgentsAutomationCost

Routines: verified recurring work for small teams

We compared Hive against thirty agent products in September 2026. Nobody pays for orchestration anymore. They pay for one recurring job that runs correctly, gets checked, and never blows the bill. That is the feature we're building.

We spent 18 September 2026 reading the competition: canvas tools, agent frameworks, “AI employee” SaaS, agent control planes, the closed vendors, and the local chat apps. Around thirty products. We also re-read Hive’s own code — flows.py, rooms.py, llm.py, triggers.py — to see honestly what we had and what we didn’t.

The conclusion fits in one paragraph. The 2026 market does not pay for orchestration. The open-source core is free everywhere — n8n, Dify, CrewAI, LangGraph, Paperclip. What people pay for is a repeated job that runs correctly, with someone checking the output, without wrecking the invoice. Hive already had most of the parts. The work is not another subsystem; it is wiring the existing pieces into one thing a small team would actually buy. We’re calling it a Routine.

What the market is complaining about

Read enough G2, Reddit and Hacker News threads and the same four complaints keep surfacing.

Cost surprises. Zapier sits at 1.4 stars on Trustpilot, largely over billing (see the Zapier Agents complaints). Lindy’s credits are opaque to users. Make charges 50 credits per agent run. Gartner’s token-usage report describes the same surge from the buyer’s side.

Brittle agents. CrewAI users ask how to stop infinite loops eating tokens. n8n’s AI Agent node has a memory failure thread long enough to be its own documentation. LangGraph shipped a self-hosted flaw chain in June.

Nobody checks the output. This is the quiet one. Almost every product runs the agent and hands you the result. Human-in-the-loop, where it exists, is a gate on a tool call (n8n 2.6’s HITL tools), not a second model reading the deliverable and saying whether the numbers are right.

Self-hosting is hard. Dify and LibreChat mean a stack of containers (a six-month Dify review is blunt about it). Open WebUI changed its license. OpenClaw exposed 21k instances to the internet.

Two more signals told us where the floor and ceiling are. OpenAI’s AgentKit is shutting down on 30 November 2026: a closed orchestrator from the biggest vendor in the space did not survive a year. And Paperclip went to 30k GitHub stars in three weeks on one idea: agents organised like an org chart, each with a budget that pauses the agent at 100 %. People are not starving for another framework. They want to know the bill will stop.

On price, Claude Cowork, Claude Code routines and Managed Agents set the frame: a routine is a recurring job with a deliverable, and 20–30 USD per seat is roughly what people will tolerate. Above that they self-host.

What the research says about “more agents”

There is a second thread, from the papers rather than the forums. The Cost of Consensus (arXiv, May 2026) and the earlier problem-drift work both find that free, homogeneous debate between models performs worse than isolated self-correction — the agents drift off the problem together. The pattern that wins is a verifier triggered on condition: one producer, one checker, and a heavier process only when the checker objects.

That matters to us because Hive has both shapes. Rooms run a /goal loop (producer, then checker, two rounds) and a full council (four phases, 2 + 2N model calls for N seats, up to six seats). The research says: make the first one the default, and stop treating the second one as the grown-up option.

What Hive already had, and the gaps

Being honest with the code, the pieces were there:

And the gaps, which turned out to be plumbing rather than architecture:

The Routine

A Routine is one recurring job, end to end, under a budget that stops itself:

trigger → room /goal → producer on a cheap or local model → checker on a strong model → conditional escalation → deliverable with signoff → delivered to the team’s channel

The customer-shaped version: “every morning at 8, collect Linear tickets and GitHub PRs, write a one-page brief, have the checker verify the numbers, post it to the team Telegram.” Or: “order form webhook → extract fields → match against the price list → export xlsx → wait for accounting to approve.”

Step by step:

  1. Trigger. Cron, HMAC webhook, email/ICS poller, and later Telegram in. The payload is attached to the run as an input with a SHA-256, which attach_input already does.
  2. Produce on the cheap tier. The producer runs on the lowest rung of the model ladder, preferring a local endpoint. route_down, today an opt-in that only applies to the first attempt, becomes the default for Routines.
  3. Verify on condition. The checker runs on a strong model in task prompt mode and returns a JSON verdict: pass | revise | escalate, a list of faults with quotations, and a confidence. On revise, the producer receives only the verdict and the rejected passage, not the whole draft — roughly half the tokens of the current second round. Two rounds maximum, which is already GOAL_ROUNDS. On escalate, or two rejections, two arbiters who were neither producer nor checker rule on exactly the disputed criteria — blind, in parallel, one call each. A defect clears only when both mark it ok with evidence; the merge is deterministic, so there is no extra Lead call. No cross-examination.
  4. Hand over. The deliverable is a file in the room, signed off against a manifest, with docgen if the output is a document. The human gate is the existing 300-second tool approver on Today, later a Telegram approve button.
  5. Deliver. Telegram out (exists), plus an outbound webhook — POST url {run, status, files, cost} — so the customer’s n8n, Make, Zapier or Google Sheet receives it, and the OpenAI-compatible /v1 API so their existing tools can call Hive.
  6. Budget that pauses. Each Routine carries a monthly budget_usd. Hitting the ceiling pauses the Routine, notifies Telegram, and makes no model call. This reuses the ledger that already reserves and settles spend.

The cost comparison

This is the whole argument in numbers.

A full council with six seats costs 2 + 2N = 14 model calls on a strong model, every time, regardless of how hard the question was. A Routine’s expected path is 2–3 calls: one producer call on a cheap or local model, one checker call on a strong model, and occasionally one revision. The worst case — two revisions and an escalation to the two-seat mini-council — is about 7 calls. The heavy lifting happens on the cheap tier; the strong model reads, it doesn’t write.

Around that, the token work, in order of payoff per effort:

#ChangeEffect
1Keep the stable prefix (harness, skills, catalog) ahead of the volatile board so gateway prompt caching keeps hitting60–90 % of system input tokens on every room turn
2Drop _no_cache on delegate sub-calls (the key already covers the specialist’s own prompt)Cache hits on repeated tasks
3Structured verdict + a fix round that lists only the failed criteriaAbout half the tokens on /goal’s second round
4Parallel delegate fan-out (a thread pool, like the room “wave”)N× faster delegate workflows
5A local rung on the ladder; route_down covers the /goal producer’s first draftProducer close to 0 USD
6Council rounds=1: blind positions → verdict, no cross-examinationCouncil down to 2 + N calls

The rule we’ll document: local for production, extraction and classification; cloud for judgment and verification. Hive’s Brain (smart engine, graph, Origin) still needs Claude and stays that way. Routines don’t go through the Brain.

What we are not building

The market told us this as clearly as it told us what to build.

No drag-and-drop canvas — n8n, Dify and Flowise have won that, and Hive composes work with Rooms, not nodes. No fourth orchestrator next to CrewAI, LangGraph and Microsoft’s Agent Framework (AutoGen is already in maintenance mode). No seventy connectors. No agent message bus or presence. No council as the default. No local Brain. No marketplace, seat billing or multi-tenancy. A2A — 150 organisations and counting — waits until a customer asks for it. Wrapping Hive as an MCP server, so Claude Code, Cursor or n8n can call a Routine as a tool, comes first.

Sequence

Three waves, each locked by an eval before the next starts.

  1. Local presets (Ollama, LM Studio, vLLM) with a health check; a local rung on the ladder; triggers that target room:/goal; outbound webhook. Acceptance: a real cron runs /goal on Ollama, the checker on cloud, a file is produced, the webhook receives the payload.
  2. Structured verdict and a fix round scoped to the failed criteria; arbitration on dispute; parallel delegates; cache on delegate sub-calls. Acceptance: token counts before and after on five eval cases; cache hits > 0 on delegate; a two-seat council in ≤ 5 calls.
  3. Telegram in with an approve button; per-Routine budget pause and notification; a Usage page per Routine; a landing page that sells Routines rather than seats. Acceptance: hit the ceiling → paused, zero model calls; a restore drill passes.

Why this fits Hive

Hive is a single process of standard-library Python and SQLite, self-hosted on your own box, with a Preact frontend that needs no build step. That is the answer to the fourth complaint before we write a line of Routine code: one process to back up, one file to restore, no container stack. The other three — cost, brittleness, verification — are what a Routine is for: a bounded job, a cheap producer, a strong checker that only escalates when it disagrees, and a budget that pauses itself instead of surprising you.

The market is paying for operational convenience and governance, not orchestration IP. That is where we’re pointing Hive.

Want something like this built for your team?

Get a quote →