We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

Every AI coding tool we'd used had the same failure mode: one model, one long unbroken stream of output, no one checking its work. It writes fast and confidently, and it's wrong in exactly the ways a junior engineer working alone at 2am is wrong, nobody in the room to say wait, that file's already correct, why are you rewriting it or you can't touch that, I own it.

Real software doesn't get built by one person typing alone. It gets built by a team that argues about scope, splits ownership, reviews each other's pull requests, and occasionally gets stuck and has to escalate. So we asked the obvious, slightly absurd question: what if the team itself was the product? Not a chatbot with a system prompt that says "you are a helpful coding assistant", an actual company that comes into existence for exactly as long as your mission takes, then answers your follow-up requests afterward like it never left.

It's not an AI agent. It's an AI engineering company that builds itself for every mission.

That sentence became the spec.

What it does

You give Orvix one sentence — "Build a weather dashboard using the OpenWeather API with search, favorites, loading states, error handling, and responsive design" is the one we ran obsessively and it:

  1. Plans without you watching. POST /missions returns in under a second; in the background, MasterMind analyzes the mission and drafts the Orvix Map — a locked contract listing every screen, file, and acceptance check the finished product has to satisfy.
  2. Designs its own org chart. Strategy Weaver reads that map and decides, for this mission specifically, how many specialists it needs and what each one owns — a scaffolder, a data modeler, a hook engineer, one feature agent per component, an integration agent. Nobody templated this roster in advance.
  3. Builds in parallel, for real. Each agent gets its own git worktree branch and holds a genuine multi-turn tool-use conversation with Qwen — reading files, writing code, committing, opening a pull request — not one blind generation pass.
  4. Coordinates without a shared brain. Agents don't dump everything into one context window. They post to the Orvix Book, a shared ledger of questions, contracts, and decisions, each agent seeing only the filtered slice relevant to it.
  5. Gets reviewed by a second AI. Critic Council checks every PR against the Orvix Map — using the full current content of every changed file, not just the diff, because a file that's already correct on main produces a tiny diff and would otherwise read as "missing work."
  6. Ships only what actually runs. A runtime acceptance gate builds the real project, smoke-tests the declared routes, and asks a Qwen judge whether the shipped product matches the mission — before the mission is allowed to complete.
  7. Keeps taking requests after "done." Through an owner channel, you can steer a finished mission — "add a dark mode" — and MasterMind will route it to the right specialist, or hire a brand-new one on the spot if nobody on the team owns it yet.

It runs live on Alibaba Cloud, on Qwen through Qwen Cloud, with a terminal cockpit (Ink/React) that streams every planning stage, every agent's tool calls, and every review comment as they happen — not a progress bar, an actual transcript.

How we built it

Two runtimes, no queue, no database: apps/api (Node.js/TypeScript) holds the entire mission as one in-memory object plus a directory of real git worktrees on disk; apps/cli (Ink/React) is a pure client of it over REST and Server-Sent Events, which means the cockpit can run on a laptop while the mission runs entirely on an Alibaba Cloud ECS instance.

The scheduler is a continuous work pool, not fixed waves — revisions, Book-signal responses, PR reviews, fresh executions, and the post-merge build gate all run concurrently, refilling as jobs complete. A background pass wakes any task stuck blocked every single scheduling round, not only when the whole pool goes idle, so one agent's slow revision cycle can never quietly starve a completely unrelated stuck agent forever.

The packet-assignment problem — deciding which of up to 20 work packets in the Orvix Map belongs to which agent — turned into its own small research problem. The scoring function that resolves it now looks roughly like:

$$ \text{score}(p) = 5\cdot\mathbb{1}[\text{id}(p)\in H] \;+\; \sum_{f\,\in\,\text{owns}(p)} 2\cdot\mathbb{1}[f\in H] \;+\; \sum_{f\,\in\,\text{files}(p)} \Big(3\cdot\mathbb{1}[\text{base}(f)\in H] + 20\cdot\mathbb{1}[f\in\text{title}]\Big) $$

where $H$ is the agent's own identity and task text. That last term — a large bonus when a packet's exact file path is literally quoted in the task's own title — is the one that survived contact with a live mission; the rest is described below.

Challenges we ran into

This is the part we're most honest about, because most of it wasn't found in a test suite — it was found on live missions, on the Alibaba Cloud instance, sometimes mid-recording of our own demo video.

  • The first-match packet bug. Early on, a mission deadlocked four agents simultaneously: several UI-flavored agents all matched the same generic packet by first-match logic, then each got blocked writing files another agent legitimately owned. Fixed by moving to the scored assignment above.
  • The reviewer that couldn't see its own history. Critic Council judged pull requests from the diff alone. A file that was already correct on main produced a near-empty diff and read as "missing" — one PR got rejected seven times while the agent correctly insisted it had already done the work. Reviews now receive the full current content of every changed file as ground truth.
  • The wake-up pass that only woke up once everyone else was asleep. MasterMind's rescue pass for stuck agents originally only ran once the whole work pool drained. One agent stuck in an endless revision loop kept the pool permanently busy — so a second, genuinely blocked agent went zero wake events across 240 mission events. Fixed by running the pass every round, unconditionally.
  • Two build gates, one node_modules. The incremental post-merge build gate and the final runtime acceptance gate could both fire npm install/npm run build around the same merge wave, corrupting the build mid-write and producing a phantom "cannot find module" error for a dependency that was correctly declared. Fixed with a per-repo lock and one automatic retry.
  • The one we found live, on camera. Recording our own demo video, an integration agent — the one whose job is literally to describe, by name, every other component it wires together — got mis-scored into a different agent's packet, because its own task title legitimately reused vocabulary ("error", "loading") that also happened to be other packets' identifiers. It escalated to MasterMind five times; MasterMind correctly replied, five times, that it had no lever to fix a mis-scored assignment from inside the Book. We paused recording, diagnosed it from the mission's own live event log over the API, added the literal-file-path-bonus term above, verified it against every agent in that exact mission with a standalone harness (zero regressions across all nine), and pushed the fix.
  • Ink doesn't clear the terminal for you. Two separate rendering bugs — stale rows left pinned above a shorter new frame, and a genuinely blank screen until the window was resized (because SSH sessions can hand a remote shell a stale terminal size that never self-corrects) — cost us real debugging time before we traced the second one to Ink computing layout against wrong dimensions, fixed by probing the terminal's real size directly via a cursor-position report instead of trusting what the OS reported.

Accomplishments that we're proud of

  • A mission that finished with 13 agents, 12 pull requests, 12 approved, zero manual intervention, after we fixed the deadlock that stopped the first attempt at it.
  • A mission that survived three separate API process restarts mid-flight and resumed exactly where it left off, because state is persisted continuously, not held only in memory.
  • Finding and fixing a real coordination bug live, on a running mission, on Alibaba Cloud, and watching the same mission recover afterward instead of restarting from scratch.
  • An owner channel that doesn't just accept feedback — it can hire a brand-new specialist mid-mission if your request doesn't fit anyone already on the team, and republish the team roster automatically so everyone else finds out too.
  • Refusing to fake success anywhere: a failed Qwen call becomes a visible degraded stage or an open question in the Book, never a canned fallback dressed up as a real answer.

What we learned

The hard part of multi-agent systems isn't the agents. It's the scheduling and the information routing between them — deciding who's allowed to touch what, what a stuck agent should do about it, and what a reviewer is actually allowed to trust as ground truth. Every real bug we hit was in that layer, not in a prompt. We also learned that a rescue mechanism only works if it's unconditional — "wake up stuck agents when the queue is idle" sounds reasonable and quietly fails the moment something else is never idle. And we learned that deploying to a real cloud instance over SSH surfaces an entirely different species of bug — terminal size negotiation, ConPTY quirks — that a local npm run dev will never show you.

What's next for Orvix

Bigger missions with real independent surface area multi-service backends, infra-plus-frontend builds where the agent society's parallelism actually outruns a solo agent instead of just matching it on coordination overhead. Tighter Orvix Map generation to catch duplicate or malformed packet claims before they ever reach an agent, instead of after. And a longer-lived version of the Book and the hired roster that can carry across related missions for the same product, so Orvix stops starting from zero every time you come back to it.

Alibaba Cloud Deployment Proof

A separate proof-of-deployment video showing Orvix running on an Alibaba Cloud ECS instance

-> https://youtu.be/CaVT8MNpp8E

Built With

Share this project:

Updates

Submission history