RouteKit Shell Agentified Workflow Deep Dive
Technical deep dive on how RKS evolved from tool-level orchestration to a multi-tier agent architecture. Covers the Dispatcher/Governor/Agent hierarchy, the Product Owner → QA → Build pipeline, context window conservation, and observations from production UAT runs.
Executive Summary
In January 2026, we published a deep dive on RKS workflow orchestration — how hooks, governance tools, telemetry, and RAG-backed planning gave AI agents a structured path from backlog story to shipped code. That post described tool-level orchestration: a single agent calling MCP tools in sequence under hook enforcement.
This post describes what happened next. RKS has been agentified — the single-agent-calls-tools model has been replaced by a multi-tier architecture where a Dispatcher classifies user intent, launches specialized Governors as autonomous subagents, and each Governor orchestrates domain-specific Agents to do the actual work. The tools from the January post still exist, but they're now called by Governors, not humans. The human's role has shifted from directing tools to supervising agents.
This post covers the architecture of that evolution: the Dispatcher/Governor/Agent hierarchy, the Product Owner → QA → Build pipeline that replaced manual story sequencing, context window conservation, and observations from running the system in production.
1. The January Architecture: Tool-Level Orchestration
The January system worked like this:
The agent (or human) was the orchestrator. It decided which tool to call next, interpreted results, and handled errors. The tools enforced guardrails, but the sequencing lived in the caller's head.
What worked:
- Hooks and preflight checks caught bad operations
- Telemetry gave full traceability
- RAG grounded plans in project knowledge
What didn't scale:
- The caller needed to know the exact tool chain
- Error recovery was ad-hoc — the agent had to figure out what to do when
rks_planfailed preflight - Multi-story work required manual sequencing
- No separation between "decide what to build" and "build it"
2. The Agentified Architecture
The current system adds three layers above the tool chain:
Layer 1: Dispatcher (CLAUDE.md) — Classifies user intent. Launches Governors.
Layer 2: Governors — Autonomous subagents with strict tool chains. Each Governor type has a prompt template and a defined sequence of MCP calls.
Layer 3: Agents — Specialized workers (research, git, telemetry, dendron, delivery, recovery) that Governors delegate to for domain-specific tasks.
The MCP tools from January are unchanged. What changed is who calls them and how.
3. The Dispatcher
The Dispatcher is the entry point — defined in CLAUDE.md, it runs in the user's Claude Code session. Its job is classification and delegation:
| User Intent | Governor | Max Turns |
|---|---|---|
| "Design this" / "Research that" | Design-Research | 10 |
| Describes work (no existing story) | Product Owner → QA → Build (three-step) | 10 + 10 + 100 |
| References a draft story ID | QA → Build | 10 + 100 |
| References a ready story ID | Build | 100 |
| "Ship these changes" | Ship | 5 |
The Dispatcher reads the appropriate prompt template from .rks/prompts/, substitutes __PROJECT_ID__ and __PROBLEM_ID__, and launches a Task(subagent_type: "general-purpose").
Key design decision: The Dispatcher does NOT call MCP tools. It only launches Governors. This separation means the Dispatcher stays lightweight and the Governor owns the full tool chain lifecycle.
Multi-story orchestration
When the Product Owner Governor returns multiple story IDs, the Dispatcher first launches a QA Governor for each story to add testRequirements and advance from draft to ready. Once all stories are reviewed, the Dispatcher launches Build Governors sequentially — one per story, in dependency order. It waits for each to complete before launching the next, reporting progress between builds. This three-step pipeline (Product Owner → QA → Build) replaced both manual story sequencing and manual test planning.
Return contract
Each Governor returns a structured result that the Dispatcher interprets:
- Product Owner
review: Present story summaries. Launch QA Governor for each story. - QA
review: Story has testRequirements, phase isready. Proceed to Build. - Build
complete: Report artifacts (branch, PR, files changed). - Build
reviewwith no orphaned tests: Auto-proceed (mechanical decomposition). Launch QA for each child, then Build. - Build
reviewwith orphaned tests: Stop for user review (scope change). - Build
failedwithtestsFailed: true: Report diagnostics — partial diff, refinement suggestions, retry count. Do not auto-retry at Dispatcher level. - Any
failed: Report error, suggest telemetry diagnostics.
4. Governors: Autonomous Subagents
Each Governor is a subagent launched with a prompt template that defines its exact tool chain. The Governor calls MCP tools in sequence, passes a session token for authorization, and returns a structured result.
4.1 Governor Session Lifecycle
Every Governor session starts with rks_governor_init:
rks_governor_init({ projectId, problemId })
→ { ok, token, flowType, branch, baseBranch, message }
This call:
- Creates or resumes a session
- Issues a token that authorizes subsequent MCP calls
- For story flows: validates the current branch and auto-checkouts to base branch if needed
- Returns branch state so the Governor knows where it's working
The branch validation at init is critical — without it, the entire refine/research/plan cycle can run on the wrong branch, wasting turns before rks_plan preflight catches the mismatch.
4.2 Build Governor Chain
The Build Governor follows a strict 6-step chain:
Step 0 — Init: Session bootstrap + branch validation
Step 1 — Refine: Analyze story quality. If suggestions returned, apply them and re-refine (max 3 iterations). If decomposition triggered, STOP and return to Dispatcher.
Steps 2-3 — Research + Re-refine: Research agent explores the codebase for target file implementations. Results feed back into a second refine pass with codebase context.
Step 4 — Plan: Background plan generation with RAG-grounded snippets.
Step 5 — Plan Review: Poll until plan generation completes. This is mandatory — the server blocks all other tools during planning.
Step 6 — Exec: Apply plan steps, run validations, auto-ship on success (commit → branch → PR → merge → cycle-complete).
4.3 Product Owner Governor
The Product Owner Governor creates stories from task descriptions:
governor_init → [research] → [story creation at draft phase] → return storyIds
It researches the codebase to understand scope, creates one or more backlog stories with frontmatter (targetFiles, acceptance criteria), and returns story IDs at phase draft. The Product Owner Governor intentionally does not add testRequirements or advance stories to ready — that responsibility belongs to the QA Governor, which has a separate research pass focused on test patterns and coverage.
4.4 QA Governor
The QA Governor bridges the gap between story creation and build execution. It operates in two modes:
Story review (pre-build): Reads a draft story, researches the target code and existing test patterns, generates concrete testRequirements from acceptance criteria, adds test file targets, and advances the story to ready.
governor_init(flowType: 'qa') → dendron_read_note → agent_research → update testRequirements → update targetFiles → set phase: ready
Post-build validation (on demand): Runs the project's test suite and reports structured pass/fail results. This mode is invoked when the Dispatcher needs test verification outside the normal build flow.
The QA Governor cannot call build or ship tools — its tool allowlist is restricted to research, dendron, and test execution. This separation ensures test planning is an independent review step, not a side effect of story creation or build execution.
4.5 Ship Governor
The simplest Governor — a one-shot flow:
rks_ship({ projectId, message })
Commits, creates a feature branch, opens a PR, merges, and returns to base branch.
4.6 Design-Research Governor
For non-build work — research queries, architecture notes, design documents:
governor_init → agent_research → dendron_create_note → rag_embed
The note is the deliverable. No plan or exec involved.
5. The Agent Subsystem
Governors delegate specialized work to agents. Agents run server-side (no MCP round-trip) and have focused tool access:
| Agent | Purpose | Key Capabilities |
|---|---|---|
| Research | Codebase exploration | RAG queries, code search, result synthesis |
| Git | Branch/commit operations | 19+ git tools, protected branch enforcement |
| Telemetry | Event analysis | Query filtering, pattern detection, failure triage |
| Dendron | Note management | CRUD, frontmatter parsing, schema validation |
| Delivery | Full release orchestration | Composes Story + Ship + Cycle Complete agents |
| Recovery | Diagnostic and repair | Git state, locks, hooks; cross-agent delegation |
Research Agent in the Build Flow
The Research Agent is central to build quality. When the Build Governor calls rks_agent_research, the agent:
- Queries RAG for relevant documentation
- Searches the codebase for target file implementations
- Synthesizes findings into structured output
- Returns results (or escalates if it can't find what it needs)
Research escalations appear in telemetry as agent.research.escalation — useful for diagnosing when stories reference files that don't exist or have moved.
6. State Machine and Phase Enforcement
The Governor state machine (governor-state.mjs) controls which tools are allowed in which state:
init → refining → planning → planned → executing → executed → shipping → shipped
Each state defines allowed tools. Calling a tool outside its allowed state returns an authorization error. This prevents Governors from skipping steps — you can't call rks_exec without having a completed plan.
Auto-Phase Transitions
After successful tool completion, the system auto-advances the story phase:
| Operation | Transition |
|---|---|
rks_plan |
ready → planned |
rks_exec |
planned → executed |
rks_story_ship |
executed → implemented |
These transitions are emitted as auto_phase.transition telemetry events.
7. Story Decomposition
When rks_refine determines a story is too large, it triggers decomposition via rks_refine_apply:
Mechanical splits (no orphaned tests): The Dispatcher auto-proceeds, launching a QA Governor for each child story (to add testRequirements), then a Build Governor for each. No user review needed.
Scope changes (orphaned tests present): The Dispatcher stops and presents the orphaned requirements to the user. Tests that don't align with any child story indicate a scope gap that needs human judgment.
8. Telemetry: What Changed
The telemetry system from January is unchanged in structure (JSONL, daily rotation, 30-day retention), but the event vocabulary has expanded significantly:
New Event Types
| Category | Events |
|---|---|
| Governor | governor.session.created, governor.state.transition, governor.branch.corrected |
| Agents | agent.research.started, agent.research.complete, agent.research.tool_call, agent.research.escalation, agent.git.started, agent.git.failed |
| Preflight | preflight.mcp_tool, preflight.branch.failed |
| Phase | auto_phase.transition, auto_phase.invalid |
| Cycle | cycle.complete |
Monitoring a Live Run
With the expanded telemetry, you can monitor a build governor's progress in real-time:
rks_telemetry_report({ projectId, reportType: 'summary' })
Returns operation counts (plan success/fail, exec success/fail), agent invocations (completed, failed, tool calls, escalations), and overall health. Combined with rks_telemetry_query for event-level detail, this gives full visibility into what a Governor is doing without watching its terminal.
9. What Stayed the Same
The agentification was additive — the January infrastructure is still the foundation:
- Hooks still enforce policies at the tool call boundary
- Preflight checks still validate before plan and exec
- RAG still grounds plans in project knowledge
- Governance MCP tools (gov_test_run, gov_lint_check, etc.) still run validation
- Telemetry still uses the same collector/storage/query infrastructure
- The phase state machine (draft → ready → planned → executed → implemented → released) is unchanged
The tools didn't change. The orchestration above them did.
10. Evolution Summary
| Aspect | January (Tool-Level) | February (Agentified) |
|---|---|---|
| Orchestrator | Human/single agent | Dispatcher → Governor → Agent |
| Tool sequencing | Caller's responsibility | Governor prompt template |
| Error recovery | Ad-hoc | Governor returns structured failure; Dispatcher decides next step |
| Multi-story work | Manual | Dispatcher sequences Build Governors |
| Story creation | Manual backlog notes | Product Owner Governor auto-creates stories |
| Test planning | Manual or skipped | QA Governor generates testRequirements from acceptance criteria |
| Branch management | Plan preflight catches late | Governor-init validates + auto-checkouts early |
| Codebase understanding | RAG at plan time | Research Agent before plan, feeding into refine |
| Decomposition | Manual story splitting | Automated via refine-apply with orphaned test detection |
| Observability | Tool-level telemetry | Tool + Governor + Agent telemetry |
11. Context Window Conservation
Context conservation was an intentional design goal of the multi-tier architecture. Each tier — Dispatcher, Governor, Agent — operates in its own context window, and the conservation effect is significant. In practice, it's what makes large autonomous runs possible.
The problem
Every MCP tool call returns a result that stays in the caller's context window. A single Build Governor cycle (init → refine → research → refine → plan → plan-review → exec) generates 30-40 tool call results. A multi-story build multiplies that. A single-agent architecture would accumulate all of these in one context window, hitting limits well before finishing complex work.
Three-tier isolation
The Dispatcher/Governor/Agent architecture creates natural garbage collection boundaries:
| Tier | What it sees | Lifetime |
|---|---|---|
| Dispatcher | Task launch + return summary (~2 tool results per story) | User session |
| Governor | All MCP tool results for one story (~30-40 results) | Discarded when Task completes |
| Agent | Server-side tool calls (RAG, code search) | Discarded when agent returns |
When a Governor Task completes, its entire context — every refine result, every plan review poll, every exec output — is garbage collected. The Dispatcher only retains the structured summary.
Real numbers: Agent User Acceptance Testing (UAT)
A production UAT run building multiple stories across a calculator application provides concrete measurements:
| Metric | Count |
|---|---|
| Governor sessions | 18 |
| MCP tool completions (Governor-level) | 574 |
| Research agent tool calls (server-side) | 181 |
| Total tool call results generated | 755 |
| Tool call results in Dispatcher context | ~40 |
| Context traffic isolated | ~95% |
Without this architecture, 755 tool results would accumulate in a single context window. The run would have exhausted context limits long before completing — and that assumes the agent could even maintain coherent sequencing across 18 Governor sessions worth of work.
Why this matters for autonomy
Context conservation isn't just about fitting within limits — it's about quality. As context windows fill, model attention degrades. By keeping each Governor's context focused on one story's lifecycle, the model making plan/exec decisions is working with relevant context, not buried under results from three stories ago.
The Dispatcher operates the same way — it tracks high-level progress (which stories shipped, which failed) without being distracted by the details of how each story was built.
12. Lessons from UAT
Running the agentified system in UAT surfaced several patterns:
Branch state is contagious. When one Governor finishes on a feature branch and the next Governor starts, the stale branch propagates. The fix: validate and auto-checkout at governor-init, before any refine or research work.
Research escalations surface story-code mismatches early. When a Research Agent cannot find target files, it escalates. These escalations often indicate stories referencing moved or renamed files — catching this before plan generation saves an entire build cycle.
Turn budgets matter. Product Owner Governors with too few turns cannot create complex multi-story scopes. Build Governors need enough turns for the full refine → research → refine → plan → poll → exec chain. Setting these too low causes silent truncation.
Telemetry monitoring works. Watching governor.session.created, agent.research.complete, plan.prompt.assembled, exec.complete, and story_ship.success events from another session gives a reliable progress view without terminal access to the running agent.
Conclusion
The move from tool-level orchestration to an agentified architecture did not replace the January system — it wrapped it. Hooks, preflight checks, telemetry, and RAG still do the enforcement and grounding work. What changed is the coordination layer: Dispatchers classify intent, Governors own tool chains, Agents handle specialized domains, and the Product Owner → QA → Build pipeline ensures stories are created, reviewed for test coverage, and built in a consistent sequence.
The core design philosophy remains the same: guardrails, not walls. The system should make bad paths hard and good paths easy. The agentified layer moved that principle up one level of abstraction — from individual tool calls to autonomous agent workflows.
Based on observations from UAT runs including uat-agents-21 and production usage across multiple projects. Last updated: 2026-03-04 — Added QA Governor (Product Owner → QA → Build pipeline), updated Product Owner Governor to stop at draft phase, added test failure diagnostics to Dispatcher return contract, and added QA step to decomposition flow.
