Running coding-agent fleets in production and writing down what breaks. Agent-fleet operations: dispatch, verification, cost accounting, GitOps guardrails.
1/ I don't build software in a chat window. I run fleets of headless agents pulling small, well-specified tasks from a queue while I do something else. Here's the whole process: seven stages, and the failure mode behind each 👇
One thing I appreciate about my harness is it's becoming better at self-healing.
z.ai is deprecating glm-4.7 which has been my workhorse for billions of tokens.
Workers figured it out and reconfigured themselves automatically.
I didn't realize until now.
A TODO list is the wrong data structure for agent work.
A dependency graph exposes only tasks with zero blockers. Agents pull from that frontier, completed work reveals the next layer. If a task is too broad, NEEDLE splits it into children.
Agents need a queue with invariants.
There's a Pixel 6 on my Tailscale mesh whose only job is to be the fleet's failover web connection, when the server's HTTP path fails, an agent checks the URL through Chrome on the phone. Different network, honest second opinion.
1/ Self-hosted runners don't save you from a GitHub Actions outage. The queue-and-dispatch layer is still GitHub's, so a control-plane incident stalls your runners like the hosted ones. That's what pushed me off. Here's the CI I run instead, end to end 👇