One morning in Hive: the ops brief, the planning room, and the PR that waits for you
Two days after we wrote up Routines, we moved a real daily workflow into Hive: a verified AWS and GitHub brief at 08:30, a feature-planning room, a weekly dependency bump that opens a PR and waits for a human, and budgets that stop themselves. Here is what ran, what broke, and the fix — a dead model alias was killing rooms quietly.
On 18 September we wrote down what a Routine is: one recurring job, a producer on a cheap model, a checker on a strong one, arbitration only on dispute, a budget that pauses itself. Two days later we did the obvious next thing and moved a real morning — the operator’s own — into Hive, the way a small company would run it. Not a demo. The AWS account is the real account, the GitHub PRs are the real PRs, and the numbers below are what the brief said on 20 September 2026.
This post is the log of that morning: four jobs that ran, one thing that broke, and the fix. The broken part is the point, so we’ve kept it in.
08:30 — the ops brief
The first Routine is a room called “Reveal · Ops sáng” (Reveal is the product the team ships; ops sáng is Vietnamese for morning ops). It runs on a schedule at 08:30 and has three seats.
The producer runs on grok-4.6 — a cheap, fast model — with exactly two tools: aws_status
and daily_brief. It writes one file, outputs/OPS.md. The checker does not read the
producer’s prose and nod. It re-calls the same two tools itself and verifies every number in the
file against what the tools return. A lead seat exists but only speaks when the checker and
producer disagree. That is the conditional-arbitration shape from the research: no debate unless
there is a dispute.
aws_status is read-only. It pulls the CloudWatch alarm list, the RDS and EC2 inventory,
month-to-date cost by service — including the Amazon Bedrock lines, because that is where
“hidden” AI spend hides in an AWS bill — and the Budgets API. daily_brief covers GitHub: PRs
waiting for the operator’s review, team review requests, the operator’s own open PRs, and
assigned issues, across two GitHub accounts.
What the brief said this morning, and what the checker confirmed:
- AWS month-to-date: $1,521. RDS $804, EC2 $567, EC2-Other $119. The Bedrock lines total about $0.15. The checker’s verdict on the whole file was one line: “all numbers and alerts match aws_status + daily_brief exactly; files verified.” The whole run cost $0.13.
- Inventory: 8 EC2 instances, 1 Aurora PostgreSQL instance.
- 4 CloudWatch warnings. Two are RDS alarms stuck in
INSUFFICIENT_DATAand both point at a DB identifier that no longer exists. A human glancing at the console had scrolled past those for weeks. The snapshot flags it asrds_target_missing; the producer listed it and the checker confirmed the alert list matched. That one line is the clearest return on the checker we’ve had so far. - GitHub: 1 PR waiting for the operator’s review (a Dependabot
flaskbump). 12 open PRs of the operator’s own, spread acrosshivedashboardandampd-reveal. Twelve is too many; the brief now says so every morning until it isn’t.
The brief also reports on Hive’s own bill. AI spend in Hive this month is about $134. Of that, $125 is a worst-case estimate for 25 legacy runs that predate per-run pricing — we price unpriced runs at the ceiling rather than at zero, so the number errs high on purpose. The caps are unchanged: $20 per day for the workspace, $5 per run.
The planning room
The second job is not on a schedule. “Reveal · Kế hoạch tính năng” (feature planning) is a standing room with four seats: a Product Lead who leads, a Tech Lead and a Cursor Engineer who do the work, and Claude Opus via the Brain as the checker.
You type /plan. The lead plans, the workers run in parallel waves, the checker reviews, and
the room writes outputs/REPORT.md with one row per feature: feature · value · effort · risk
· owner · week. It is deliberately a table and not an essay. The value of the room is that
the operator reads a table over coffee instead of running three chat sessions and merging them
by hand.
This morning’s run: the Tech Lead assessed 8 features (12–18 person-days), the Cursor Engineer marked 6 of 8 as automatable by a Cursor Cloud agent, and Claude’s critique cut the scope to 5 features (about 13 person-days including review), deferring the Lambda and multi-model work. The whole room cost $0.59.
The PR that waits for you
The third job is new today: an Automation of kind “Origin task.” Origin is Hive’s autonomous coding agent — it clones a repo, edits, runs tests, opens a PR. Until this morning you launched it by hand. Now it can be scheduled.
The weekly job runs against ampd-reveal. It checks whether pipecat-ai and google-genai
have newer releases, bumps them if so, runs the test suite, and opens a PR. The configuration is
three header lines at the top of the prompt:
repo: https://github.com/luonghongthuan/ampd-reveal
base: main
approval: push
approval: push is the whole point. The run does everything up to the push and then stops.
Nothing lands without a human clicking approve. We did not build a new gate; Origin already had
plan and push approvals. The scheduled task just defaults to the push gate. A dependency PR opened by an agent on a Sunday and waiting for Monday’s
click is exactly the size of autonomy a small team should be comfortable with.
Budgets that stop themselves
The fourth piece is the least visible and the one we would sell first. Every layer of this morning’s work has a ceiling:
- The Routine has a monthly budget.
- The room has a budget.
- Each agent is capped at $5.
- The workspace is capped at $20 per day.
Hitting any ceiling pauses the thing that hit it and makes no model call. The ledger reserves before a call and settles after it, so the pause happens before the spend, not after the invoice. This is the Paperclip lesson from the research — the agent pauses at 100 % — applied at four heights instead of one.
What broke
Now the honest part.
Yesterday the chief-of-staff rooms started failing with “Model error 502.” Nothing in the
logs pointed at a room; the rooms themselves looked fine. The root cause was a fallback. When a
seat had no model set, the code picked the first model in the pool list. The first model in
the pool list was xp-claude-opus-5, an alias on a prepaid gateway whose upstream now answers
upstream_unsupported. Every room that relied on the default was quietly routing to a model
that no longer existed, and the 502 was the gateway’s way of saying so.
We fixed it in three moves:
- One
default_model()function that readsHIVE_DEFAULT_MODEL, falling back to the pool list. It replaced every hardcodedgpt-5.5fallback inflows.py,rooms.py,evals.pyandroutes.py. There were more of those than we’d like to admit. - The pool list reordered,
grok-4.6first, because the default should be the cheap model that works, not the expensive one that might. - Dead aliases removed: the
xp-*family,gpt-5.5,gemini-2.5-pro(removed upstream), andgrok-3-mini-fast(out of credits).
The audit turned up more than the 502:
gpt-5.5andcdx-gpt-5.5through the gateway return an HTML login page instead of JSON. The codex route on the gateway needs a re-login. Hive now routes around it rather than parsing a login form as a completion.deepseek-chatanswers 402, insufficient balance.- The Linear MCP token has expired and needs a reconnect in Apps.
- Google Calendar isn’t connected yet. It’s one OAuth click; it hasn’t been clicked.
- Telegram delivery of the daily brief needs the operator’s chat id on their profile. Until then the brief is a file in the room, which is fine, but not a notification.
None of these are bugs in the agents. They are the boring failure modes of running real integrations on real accounts: a token expires, a balance runs out, a vendor removes a model, a gateway route needs a login. The agents did not fail loudly on any of them. That is the lesson.
The lesson
Agents die quietly when a model alias dies. A room doesn’t crash; it returns a 502 and moves on, and if nobody is reading the room it looks like nothing happened. The fix is not better error messages. It is structural: the pool list must be the single source of truth for which models exist, and every fallback must read it. A hardcoded model name anywhere in the code is a future outage with a date we don’t know yet.
This is also why the checker earns its seat. The producer on the ops brief could have been
routed to a dead alias too. If it had, the checker — on a different model, re-calling the tools
itself — would have refused to sign off on an empty OPS.md, and the lead would have been
woken up. Two models on two routes is cheap insurance against one route going dark.
Where this leaves the product
We said on the 18th that we would not build a canvas or a fourth orchestrator, and that the thing to sell is one recurring, verified job: producer on a cheap model, checker on a strong one, arbitration only on dispute, a budget that pauses itself. Two days of running our own morning through it hasn’t changed that. It has made the pitch concrete:
Every morning at 08:30, Hive reads your AWS account and your GitHub, writes a one-page brief, has a second model verify every number, and tells you about the two alarms pointing at a database that doesn’t exist. Once a week it bumps your dependencies and opens a PR that waits for your click. It cannot spend more than you told it to.
That is the job. Everything we fixed today was in service of making it run tomorrow without anyone watching.
Want something like this built for your team?
Get a quote →