Running work from the plan
Run an item or a document — the checklist, Advanced settings, the teams, and what moves a plan item.
The plan is the brief. A plan item already carries a six-section body written before the work started, so there is no box to type a brief into on the way to running one — the run button takes the plan you are already looking at and queues it.
Where the button is
| Where | Button | What it queues |
|---|---|---|
| The item page | Run this item | That one item |
| A document | Run this document | A research run against that document |
| The Floor, when the project has a probe target | Run probe | A one-implementer probe run |
There is no phase-wide or release-wide run button on the plan page in this release: a run targets one item, one document, or the probe. A run can still carry several items when a launcher hands it a checklist — the popover below draws one whenever it has more than one item to offer.
It opens a popover, not a dialog — taking the plan away to show it back to you is the wrong move.
The checklist and the summary
When the popover has a list of items, each gets a checkbox, pre-checked for every item that is not done, not already in a live run, and has a complete brief. An item with no complete brief cannot be checked at all; its row says needs a complete brief — and what is missing. At most 200 items go in one run, and a larger selection is refused in plain words rather than trimmed. A single item or a document gets the same popover with no list, because there is only one thing — and if that one item has no complete brief, the popover says so and Queue stays off.
Under it is one derived summary of what Queue is about to do:
Queues 3 runs with the build and verify on vm-1.
Environment: test. …It is built from the current settings rather than written down, so it cannot describe something other than what the button does. With no machine chosen it ends on any online machine — but Queue still asks you to choose one: "Choose the machine this should run on." The machine is named, never inferred.
The button reads Queue 3 runs. A run queued for a machine that is offline waits; nothing is lost.
Advanced settings
Everything else folds behind an Advanced settings disclosure.
| Field | Default | Notes |
|---|---|---|
| Team | From the target — item → build and verify; document → research; the Floor's probe → probe | A team no online machine holding the repository advertises is greyed — no machine knows it — unless no machine has advertised any |
| role model, one per role on the team | The first Claude model the capable online machines offer, else the first model any of them offers | Built from the machines' advertised providers; a model no capable machine offers is listed and greyed — no machine here offers it — so you learn why it is unavailable instead of wondering where it went. With nothing advertised the list says No machine has advertised a model |
| Environment profile | test | Shown only when the project has more than one profile. The secrets and manifest profile the run gets |
| Machine | Choose a machine — required | Every registered machine with its liveness. One without the repository is greyed and says so — does not hold owner/repo — as is one that cannot reach it; one that can read but not push says read-only on owner/repo |
| Per-run time limit (seconds) | Empty — the machine's own limit (3600 s) | |
| Extra instructions | Empty | Appended to the item bodies. The placeholder says it: "Optional. The items' own bodies are the brief." |
There is no effort control and no task-type control. The task type is derived from the team — "who runs" already answers "what kind of run this is", and the two were free to disagree for as long as both were inputs.
The teams
A team is a deterministic sequence over one item, not a conversation between agents.
| Team | Who runs | Ends at |
|---|---|---|
| Solo | implementer → QA | A branch, a pull request if the agent opened one, and the four suite counts |
| Build and verify | implementer → reviewer → fix loop → QA | The same, with a reviewer verdict; the default for a single item |
| Phase team | build-and-verify per selected item, in parallel up to the machine's capacity | One worktree and one branch per item |
| Plan | planner → reviewer | Phases and items written back with upsert_plan, each marked needs manual check or not from its test notes |
| Research | researcher → reviewer | One draft document with its sources |
| Custom | implementer → reviewer | |
| Probe | implementer | The Floor's Run probe |
Two rules hold across all of them:
- The fix loop runs at most twice. A reviewer verdict of
NOT_CLEANre-runs the implementer with the defects verbatim; a failing QA step feeds the same loop with the failure tail attached. After two passes the run fails with what it failed on. - A reviewer verdict, a plan and a research report arrive as structured output, never as a prose block an agent hoped would be parsed. A missing or invalid report fails the run closed rather than being read optimistically.
QA is a step, not an agent. It is a deterministic daemon step with no model in it: the repository's own check and test commands from kairoku.json, the four counts parsed off the runner's summary, the failing tail attached as the defect. That is why Advanced settings asks for a model per role and never for a QA model — there would be nothing for it to drive.
What moves a plan item
The app writes exactly one automatic item transition, and nothing else about a plan moves on its own:
not_started → in_progress, on the first running report for that item. Once. A second report never writes again, and an item you have marked blocked is left alone.
done is always a person's. A merged pull request is evidence that the code landed, not that anybody signed off, so a merge marks the run merged and leaves the item where it is. Merging is yours too — the run page says so in as many words, merging stays yours, beside the PR link. Merge detection is a read against GitHub, with your GitHub connection, when the Floor loads — up to five finished runs per load that have a PR and no recorded merge, in a project whose repository is on GitHub. A GitLab merge is not detected yet.
Needs manual check
needs manual check is a flag on the plan item, set by upsert_plan at plan time (the planner sets it from the item's test notes) or by hand on the plan page. Once its pull request is seen merged, an item carrying it lands on the Floor's Needs-you rail as Merged — needs your check, with Mark done beside it.
Agents cannot set done
The MCP tool update_item_status refuses done from an agent credential and answers with a reason. in_progress and blocked are still an agent's to write. Done comes from a person — see the plugin's protocol table.
A release's stage never moves automatically. Shipping is a gate, not a status.
After a run
Run again appears on a failed run, on the Floor and on the item pane. It queues a fresh dispatch for the same item with the same team and models, on a new branch, with the previous pull request's link and the failure summary appended to the brief — so the second attempt starts knowing how the first ended.
Operator
An operator chat can inspect an authorized release and prepare an exact action. Confirming that action is a separate human request. Closing the chat does not cancel a release that was already admitted, and operator text cannot merge or mark an item done.