When the model can drive itself: we rebuilt our navigation framework and open-sourced flightwake
Newer models no longer need step-by-step navigation, but four problems remain. Why we rebuilt our AI workflow framework as a dashcam and open-sourced it.
Kai Wu
• Founder, Kaiwu TechEngineeringPublished Jul 19, 20267 min read
For the past two years, the standard way to get AI to write code has been "navigation": research → plan → execute → verify, one gate at a time, with a checkpoint at every step. We built a stage-driven navigation framework of our own along those lines and used it every day. It is what carried us through 12 projects and more than 2,000 commits in two months (in Chinese).
Then the Fable 5 generation of models arrived, and we noticed something awkward: the model no longer needs navigation. It can drive on its own.
Give it a clear goal and it explores, breaks the work down and verifies the result by itself. The planning ability our framework used to supply is now fighting the model for the steering wheel, and the framework is often the worse driver.
So we made a decision: tear down the navigation framework we had used for two years, rewrite it as the opposite kind of tool, and open-source it. It is called flightwake. It is not a navigator. It is a dashcam: a flight recorder for work done with AI agents.
Four things even a stronger model cannot do
Dropping navigation does not mean dropping everything. We went through two years of our own collaboration records and found four problems that are structural. They will not go away as models get better:
- Sessions end, and context is finite. Every chat window gets closed eventually. Without a record, the next session has to start with an archaeology dig through git history.
- Git records what, not why. A commit tells you what changed. It does not tell you why the other path was rejected, or what actually caused that pitfall.
- Discipline drifts in long sessions. As work stretches out, even the strongest model will sometimes say "done" before the tests have run.
- Claude, Codex, Gemini and humans do not share state. Unless the state is in git, every agent and every person is working in their own parallel world.
What the four have in common is that the missing piece is not intelligence. It is persistence and discipline. The navigation framework was fixing the wrong thing.
What is left after navigation: three parts
flightwake's core principle fits in one sentence: records follow the work; they do not direct it. By default you just start working, and the framework stays out of the way. It records from the side, like a dashcam. There are only three components:
The dashcam: three append-only files. DECISIONS.md (every decision that closes off other options, one line each, including the why), TRAPS.md (non-obvious pitfalls: symptom, root cause, workaround) and records/ (a flight record written as each piece of work wraps up). All plain Markdown, all committed to git.
Warning lights: a few hard guardrails that do not depend on how capable the model is. Work counts as done only when tests pass and typecheck is clean. Any change that touches production must leave verification evidence in the record. Destructive operations need the user's confirmation first.
Road signs: a STATE.md that always reflects the current situation: where things stand, what is in progress and where to pick up next. Any new session can take over after reading it, whichever vendor's model it runs on, and so can a human colleague. In our own tests, going from a cold start to a safe handover takes under 2 minutes, at a marginal cost of 3–5K tokens and about 19 seconds of API time.
The cost model is also the opposite of a navigation framework's: it is event-triggered. You write a line only when you make a decision, log a trap only when you hit one, and write a record only when you wrap up. A session that triggers nothing costs zero extra tokens. A gate-based framework charges a toll on every task, large or small.
The dashboard: red means it is time to wrap up
With the optional status line installed, a gauge stays at the bottom of the terminal: health color, how many commits STATE is behind, and context usage. It does more than show status. It suggests the next command at the right moment: wrap up when context is nearly full, run a cold start when you have just opened a new session. Red means it is time to hand off. Green means the next session can pick up cleanly.
The design principle is the same as the framework's: prompts for the human, an obligations table for the model, and both are triggered passively. You do not have to memorize any rules.
A stage-by-stage playbook for working developers
Before open-sourcing, we found a real gap. For many developers, the problem is not that the model is too weak. It is that they do not know what their own job is when they work with a strong model.
So the repo includes docs/workflow.md, a stage-by-stage playbook. Each stage covers two things: what you do, and what you say to the model. It starts from this premise:
The model's capability is not the bottleneck. Your bottleneck is knowing what to do and what to say right now. You hold the steering wheel (what you want, whether it counts as done, whether it is worth continuing); the model drives (how to do it, doing it, verifying it).
The playbook splits a development cycle into six stages. Start by restoring state (do nothing yet; listen to the report first). Decide what to build (say what, not how; for anything that would hurt to get wrong, ask for a plan before giving the go-ahead). Let it work without interrupting (stop it only when the direction is wrong, and give a reason when you do). Accept only evidence (paste the test output; no verbal reports). Wrap up or hand off ("wrap up" when the work is done, "hand off" when it is not; closing the window without doing either is the most common mistake). When something breaks, stop the bleeding first (if health is not green, do not stack new work on top).
Each stage has a collapsible advanced section for experienced users, but the main path is deliberately written so that anyone who can code can follow it. The one thing in the whole playbook you actually need to memorize: start each session with /fw-coldstart. The model triggers the other obligations on its own. The human's job shifts from remembering process to making judgments. We think that is where people belong now that models are this capable.
The most honest demo: it recorded its own creation
The flightwake repo contains a .flightwake/ directory. It is not sample data. It is the project's real record, kept by flightwake itself: from the gap list, to installing itself, to preparing for open source, to publishing on npm. Every flight record in it was written by the agent as it worked. If you want to see what the framework looks like in use, read its own dashcam footage.
A few engineering commitments while we are at it: zero dependencies (Node built-in modules and git only), plain Markdown, MIT license, and npm releases through trusted publishing with SLSA provenance. An honest caveat: the templates and CLI are currently in Traditional Chinese (we are a team in Taiwan; English defaults are planned for the next release), and the sample size behind the performance numbers is still small. The full methodology is in the repo's benchmarks docs. You are welcome to run it yourself and send your data in as a PR.
npx flightwake init
What does this have to do with small and mid-sized businesses?
On the surface this is a story about a developer tool, but the problem it addresses will look familiar: critical knowledge exists only in one person's head (or one chat window). When the person leaves or the window closes, it is gone.
Handovers that require archaeology, decisions nobody remembers the reason for, "done" with no evidence behind it: these are not problems that come with AI collaboration. Every company has them. In a different setting, the same problem is why a knowledge base needs governance before it needs tools (in Chinese). flightwake is the remedy we wrote for ourselves, and open-sourcing it simply makes the prescription public. The same principles (state lives in files, decisions keep their why, acceptance requires evidence) also run through every client project we take on.
flightwake is listed among our own products. If your team works with AI and has been bitten by handover and discipline problems, try it. If you would like help bringing this discipline into your company's development or operations, talk to us.
Tags
Related posts
More notes on similar problems