Four agents in Orca: flightwake's single-writer rule
Claude Code and Codex share one Orca window, and our project manager agent started coding after /clear. The multi-agent rules we set with flightwake.
Kai Wu
• Founder, Kaiwu TechEngineeringPublished Oct 12, 20266 min read
First, the setup. On a normal day we run several agents as tabs in one Orca window (Orca is a desktop app that hosts AI coding agents in visible terminals). In the planning repo, Codex is the project manager and Claude is the technical director. In the implementation repo, Claude writes the code and Codex reviews it. Two repos, four agents, all working under flightwake's records.
Then one day the project manager ran /clear and read "Next step: fix X" in STATE. It did not assign the task to anyone. It fixed it itself.
We covered why flightwake exists, and the four structural problems it addresses, in how we rebuilt our navigation framework as a dashcam. This post is the second half: once there was more than one agent, how our rules for dividing work in Orca and flightwake took shape.
Why we insist on visible Orca tabs, not background calls
Letting one agent call another in the background looks like the least effort. We tried it and paid for it. We ran codex exec from an agent's shell, and it sat waiting for an end-of-input signal that was never going to come: 9 minutes at 0% CPU, no session opened, no error. Only then did we realize it had never started.
The deeper problem is that the lead agent gets only the reviewer's final conclusion, not the reasoning behind it, and terminal output is not kept. So the rule is: to discuss something with another agent or ask it for a review, always send the message into an Orca tab the user already has open and read the reply there. If there is no suitable tab, ask the human first. Do not fall back to the background on your own.
Rule one: one STATE, one writer
flightwake's wrap-up reminder fires in both Claude and Codex sessions. If nobody says clearly who the writer is, both sides will edit STATE at the same time. Here is how we split it:
- The agent being consulted: reads only, changes nothing. It does not write records or touch STATE; it hands its conclusions back.
- The agent that asked: writes the conclusions it adopts into its own record, and makes any code changes itself.
- Four things in every review reply: the conclusion, the evidence (file and line numbers), what was not verified, and the alternatives considered but rejected.
- Replies that affect a decision: saved verbatim as a file in the repo, for example
docs/plans/<topic>.review-<reviewer>.md.
The value of keeping reviews verbatim is clearest in flightwake's own development. While we were building a Claude Code mod, a model from another vendor read the diff and reviewed it over 4 rounds, and we saved 4 verbatim copies. The counterexamples from round one became tests: 179 passing and 31 failing before the fix, 210 passing and 0 failing after. Round two then listed 23 adjacent cases, 12 of which failed. When we designed team roles, Codex in the same folder ran two rounds of review and, by actually testing it, overturned our assumption that a role definition file is a hard constraint. All of this is still in the repo, not in some window that has since been closed.
Rule two: roles go in the instruction files, where /clear cannot erase them
Back to the project manager from the opening. The root cause was that its role existed only in the conversation, so clearing the conversation erased it. The division of labor was written in STATE, but STATE describes what is happening right now. It does not say who you are or what you are not allowed to do.
The fix relies on one fact: Claude Code reads CLAUDE.md and Codex reads AGENTS.md, and both reread them at the start of every session, including after /clear. Write each role into the matching instruction file, and within one folder the file name itself is the identity. No hook is needed. This later shipped as fw-roles, an optional add-on.
To test it, we put the same leading question to all four roles: "There is a typo on a button. Who are you, and what do you do next?" All 4 handed it off; none of them fixed it themselves. What did the work was not the job description. It was the list of prohibitions: "the moment you catch yourself thinking 'I'll just fix this while I'm here,' stop and assign it instead."
This is also where we hit a trap. The second version generated a native agent definition for every role. Only a dry run on a real team exposed the problem: the project manager could summon a writer on the spot to write code for it, bypassing the split between writing and reviewing. Once we generated definitions only for on-call roles, that team's generated files dropped from 24 to 12.
Memory lives in the repo, not in any one model
Claude and Codex do not share conversation history, and they do not need to. flightwake's memory is four files in the repo: STATE, DECISIONS, TRAPS and records. There is only one rule: whoever stops writes; whoever takes over reads.
On October 11, Codex was building flightwake's demo site in one tab while Claude published 0.17.0 to npm in another, and each wrote its own record. Claude's release targeted the commit Codex had just pushed. There was no export and no sync step.
A trap we fell into: an overview nobody used
The most expensive lesson was flightwake-tower, a cross-repo status overview and trap-lookup tool that passed acceptance testing on 21 repos. 6 weeks after we finished it, we looked back and found that neither Claude nor Codex had installed it, and it had been called zero times in real work. Orca already showed us every project every day, so the need for an overview had been met long before. We froze it and did not release it.
So the side panel in mod 0.3.0 takes the opposite approach. Instead of building a separate reporting channel, it reads Orca's tab list directly, groups Claude and Codex into cards by repo and refreshes every 10 seconds. Each row has a "jump to" action, and you can ask an idle agent to wrap up by clicking twice to confirm. The caveats, stated plainly: 500 tests pass, but "jump to", "wrap up" and Codex status detection have not been verified on a real machine, and Codex's status is inferred from the tab title.
What multi-agent work lacks is not more agents. It is one writer for every fact and a verbatim copy of every conclusion.
Three checks for a multi-agent team
- Who writes STATE? If you cannot answer, sooner or later two sessions will wrap up at the same time, each writing its own version, and a person will end up merging them by hand.
- Where does the review process live? A review that comes back as "looked at it, no problems" is no review at all. The evidence needs file and line numbers, and anything not verified has to be stated.
- After
/clear, does the agent still know what it is not allowed to do? A role written into the conversation lasts exactly as long as the conversation.
We apply the same discipline (state goes into files, decisions keep their why, reviews keep the original text) to every client project. Tell us how your team works with AI today, and we will tell you which part needs rules first.
Tags
Related posts
More notes on similar problems