← Back to Thinking

RouteKit Shell Workflow Deep Dive

Vince Mease• Technical Deep Dive• 1/22/2026
blogtechnicalproject:routekit-shellai-agentsgovernance

Technical deep dive explaining how RKS orchestrates AI agent workflows with guardrails, hooks, governance, and telemetry. Covers the core Detect → Plan → Validate → Apply pattern applicable to code, business automation, and beyond.

Executive Summary

RouteKit Shell (RKS) provides an opinionated, auditable workflow for AI agent operations. While originally built for code changes, the core patterns—Detect → Plan → Validate → Apply—apply to any domain where agents need governance: business process automation, trading workflows, document management, infrastructure operations, or data pipelines. RKS balances agent autonomy with human oversight using hooks, governance (MCP), RAG-backed context, and robust telemetry. This post explains the problem RKS addresses, the end-to-end workflow, enforcement mechanisms, automated validation, governance primitives, knowledge management, and telemetry for observability.

1. Introduction: The Agent Autonomy Problem

AI agents are capable of executing complex tasks quickly, but this capability introduces risks: unvalidated actions, hard-to-audit decisions, and changes that bypass review processes. Whether an agent is modifying code, executing trades, generating documents, or orchestrating business workflows, teams need a structured approach that preserves agent productivity while ensuring human reviewability, automated validation, and traceability.

RKS Philosophy: Guardrails, Not Walls

RKS design principle is to provide guardrails—block risky actions while suggesting safe, governed alternatives—rather than hard walls that prevent agent productivity. Hooks enforce policy and suggest MCP tools. Plan and exec flows ensure actions are deterministic and auditable. Telemetry and RAG-backed knowledge provide context and traceability.

The patterns work for any domain with appropriate MCP or API support:

  • Software development: Code changes, test execution, deployments
  • Business automation: Trading workflows, approval chains, data processing
  • Document management: Generation, review, publishing pipelines
  • Infrastructure: Configuration changes, scaling operations, incident response
  • Data pipelines: ETL orchestration, validation gates, rollback capabilities

2. The Core Workflow

The RKS phase state machine organizes work into discrete, auditable stages. Each phase adds validation and changes the allowed actions for the agent.

Phases and Transitions

  • draft: Initial work-in-progress note or task sketch. Low risk, informal edits.
  • ready: Requirements are sufficiently captured; planners can attempt to form a concrete plan.
  • planned: rks_plan has produced step-by-step actions.
  • executed: rks_exec has applied changes in an isolated context and run validations.
  • integrated: Actions merged to integration branches/systems after passing governance checks.
  • released: Release pipeline promotes changes to production.

Key Tools

  • rks_plan: Takes a backlog story or problemId and returns an actionable plan with prioritized steps. For code workflows, supports // CREATE FILE: directives for new file creation.
  • rks_plan_ready: Validates story readiness before planning—checks phase, target files, testing requirements, and CREATE FILE directives.
  • rks_exec: Applies plan steps, manages branches/state, runs validations, and coordinates retries/self-healing.
  • rks_refine: Surfaces human-guided next steps when execution fails after exhausting retries.
  • rks_story_ship: Atomic story shipping—wraps staging_pr + staging_merge + mark_implemented + cycle_complete in one call.
  • rks_release: Creates a release cadence, bumps versions, and triggers CI/CD workflows.
  • rks_guardrails_off / rks_guardrails_on: Temporarily disable hooks for off-rail work with full session logging and automatic PR creation on restore.

Phase State Machine

Rendering diagram...

Design note: Transitions map to strict validations. rks_plan_ready gates planning with blocking checks (phase, targetFiles, testing requirements). If rks_exec fails after exhausting retries, it calls rks_refine which surfaces guidance for fixing the story before re-planning.

3. Hooks Enforcement System

Hooks intercept agent invocations and enforce policies. They follow a simple pattern: detect a risky action, block it, and suggest a managed replacement. Hooks also attach context and telemetry to every decision.

3.1 Multi-Layered Hooks Protection

The hooks directory (.routekit/hooks/) is critical infrastructure—if it's missing or corrupted, guardrails don't work. RKS implements defense in depth:

  1. Startup verification: MCP server checks hooks exist on startup. If missing, attempts auto-restore from templates/generic/.routekit/hooks/. If unrecoverable, refuses to start with clear instructions.

  2. CI protection: GitHub Actions workflow fails any PR that removes the hooks directory.

  3. Critical error messaging: rks_guardrails_off returns severity "critical" with recovery instructions if hooks are missing—can't disable what doesn't exist.

  4. Health reporting: rks_guardrails_status includes a hooksHealth object showing present/restorable/mismatch status and listing any missing or extra hooks vs. the template.

3.2 Common Enforcement Hooks

  • enforce-git-workflow.mjs: Detects raw git commands and prevents them when originating from an agent shell. Suggests rks_git_commit/rks_git_push which wrap git with validation hooks and telemetry emission.
  • enforce-read-provenance.mjs: Blocks file reads without RAG provenance to ensure agents discover information through proper knowledge channels.
  • enforce-plan-scope.mjs: Prevents direct edits on rks/* plan branches—changes must go through the plan/exec cycle.
  • enforce-branch-workflow.mjs: Enforces branch naming, prevents direct commits to protected branches, and ensures PR-based integration.

Hook Enforcement Flow

Rendering diagram...

UX implications: The hook behavior aims for low friction: suggestions include exact commands and links to required policy info. Hooks also support escape hatches for human operators under strict auditing (e.g. require an explicit acknowledge flag and emit a high-severity telemetry event).

3.3 Guardrails Off/On Governance

For legitimate off-rail work (debugging, manual fixes, experimental changes), RKS provides governed escape hatches:

rks_guardrails_off(reason)

  • Moves hooks to .routekit/hooks.bak/
  • Logs session start with timestamp, reason, branch, and HEAD commit
  • Returns session ID for tracking

rks_guardrails_on()

  • Restores hooks from backup
  • Detects changes made during session
  • Auto-ships changes via PR flow:
    1. Creates branch off-rail/{sessionId}
    2. Commits changes
    3. Opens PR
    4. Squash-merges to staging
    5. Runs cycle_complete to return to staging
  • Logs session end with duration and files changed

This ensures off-rail work is auditable and follows the same PR-based integration as governed workflows.

Guardrails Off/On Flow

Rendering diagram...

4. Automated Validation System

Validation is central to RKS—validators are the gatekeeper that determines whether an action can be committed, retried, or rolled back. For software projects, this typically means running tests. For other domains, validation might involve API health checks, data integrity verification, or external service confirmations.

4.1 Validation Detection

For code projects, the test detection routine inspects package.json and the project workspace to find the most appropriate command. Priority is:

  1. A specialized script such as test:unit when present
  2. The test script
  3. DevDependency heuristics (vitest, jest, mocha)

It returns { cmd, args } or null. This deterministic detection ensures rks_exec knows how to run validations consistently.

4.2 Validation Execution

The validation runner executes the chosen command with a configurable timeout (default 120s). It captures stdout and stderr streams, aggregates results, and writes debug traces to .rks/test-runner-debug.json. The function returns a structured result that includes pass/fail status, summary, and the raw output for diagnostics.

Result schema (conceptual):

{
  "passed": true,
  "skipped": 0,
  "output": "string",
  "summary": { "total": 10, "passed": 10, "failed": 0 },
  "exitCode": 0
}

4.3 Validation in rks_exec

The exec flow after applying changes follows a clear pattern:

  1. Apply changes from the plan into an isolated context (backup previous state)
  2. Run validations
  3. If validations pass: commit the changes, emit telemetry, and proceed to integration
  4. If validations fail: record outputs to tests-failed.log and create a summarized failure report
  5. Attempt self-healing up to MAX_RETRY_ATTEMPTS:
    • Generate a repair plan using an LLM, passing failure output as context
    • Apply the repair plan
    • Re-run validations
  6. If retries are exhausted and failures persist:
    • Roll back to the pre-change backup
    • Call rks_refine to surface human-guided next steps
    • Return a structured failure result with helpful hints and preserved logs

4.4 Retry Logic and Self-Healing

The default MAX_RETRY_ATTEMPTS is 2. Each retry is limited by scope and must pass governance checks before being applied. Repair plans are stored alongside the run results for auditability (e.g. .rks/repairs/attempt-1-plan.json).

Validation Execution Flow

Rendering diagram...

Operational notes:

  • All validation runs and retries produce artifacts in the .rks directory to support debugging and forensic analysis.
  • Fix plans are constrained by governance: they cannot modify critical resources (policy-defined) without explicit human approval.

5. Plan Readiness & Quality

Before rks_plan generates implementation steps, stories must pass readiness checks via rks_plan_ready.

5.1 Blocking Checks

  • Phase status: Must be "ready", "planned", or "executed" (not "draft")
  • Target files: frontmatter must include targetFiles array
  • Testing requirements: Must have a ## Testing Requirements section
  • CREATE FILE directives: Non-existent target files must have // CREATE FILE: path directive with content

5.2 Warning Checks

  • No test files: Warning if no test files in targetFiles
  • Multi-file stories: Warning if >3 target files (higher partial failure rate)
  • AC/step ratio: Warning if acceptance criteria count far exceeds step count

5.3 CREATE FILE Support

The planner recognizes // CREATE FILE: path directives in story bodies and generates create_file action steps with the accompanying content. This allows stories to specify exact file contents for new files.

Example story format:

// CREATE FILE: packages/foo/bar.mjs

```javascript
export function bar() {
  return "hello";
}

The planner extracts both the path and content, generating an executable create_file step.

6. Governance Tools (MCP Server)

RKS exposes governance tools via MCP (Model Context Protocol), the standard for AI tool integration. These tools provide specialized checks and long-running utilities that complement local execution.

Governance Endpoints and Responsibilities

  • gov_test_run: Runs targeted test suites with parameters testType (unit|integration|e2e|all) and watch. Useful for targeted validation before merging.
  • gov_lint_check: Runs lint across a given scope and can optionally auto-fix when fix=true.
  • gov_build_check: Ensures the project builds successfully. Detects type errors and missing imports.
  • gov_scope_validate: Validates that the set of changed files matches expectedFiles and checks for breaking changes if checkBreaking=true.
  • gov_quality_check: Runs multi-dimensional checks (docs, tests, types) and returns a quality score plus remediation hints.
  • gov_regression_check: Focused regression checks for templates, CLI commands, and design system components.

MCP is also the place where policy evolves: teams can add new checks without changing the RKS client. MCP responses include structured reasons and remediation actions which rks_exec surfaces to agents and humans.

7. Telemetry & Observability

Observability is essential to trust agent-driven operations. RKS emits structured events at every major step so teams can trace the lifecycle of actions from plan → exec → validate → commit.

7.1 Event Types

RKS emits the following logical event types (examples):

  • Planning: plan.targetfiles.parsed, plan.prompt.assembled, plan.prompt.validated, plan.rag.query_built, plan.ac.coverage, plan.reviewer_mode.start, plan.reviewer_mode.complete, plan.retry, plan.retry.exhausted
  • Shipping: ship.start, ship.step.completed, ship.success, ship.failed, story_ship.start, story_ship.step.completed, story_ship.success, story_ship.failed
  • Refinement: refine.analyze, refine.apply, refine.ready
  • LLM: mcp.llm.start, mcp.llm.complete
  • Stories: story.phase.changed
  • Guardrails: guardrails.off, guardrails.on, guardrails.auto_shipped
  • Auto-analysis: auto_analyze.start, auto_analyze.passed, auto_analyze.failed
  • Phase transitions: auto_phase.transition, auto_phase.invalid

7.2 Event Structure

Each event is a compact JSON object that carries essential metadata for tracing and aggregation. Example:

{
  "id": "uuid",
  "type": "exec.complete",
  "timestamp": "2026-01-22T15:30:00Z",
  "projectId": "routekit-shell",
  "correlationId": "uuid",
  "runId": "rks-run-abc123",
  "payload": { "stepsApplied": 5, "latencyMs": 1234 },
  "context": { "branch": "feature/foo", "actor": "agent/auto-123" }
}

7.3 Correlation IDs

Correlation IDs link related events across systems: planner → exec → validation → commit. They are generated at run start using collector.startCorrelation() and propagated to downstream helpers. Correlation IDs are crucial to reconstruct the end-to-end timeline of an agent's decision path.

7.4 Storage and Retention

Events are persisted as JSONL files under .rks/telemetry/ with daily rotation, e.g. .rks/telemetry/events-2026-01-22.jsonl. Default retention is 30 days and can be configured via RKS_TELEMETRY_RETENTION_DAYS. The collector buffers events in memory and auto-flushes when the buffer reaches a threshold (default 100 events) or when a run completes.

7.5 Where Events Are Emitted

Instrumentation points include:

  • planner.mjs: emits plan.targetfiles.*, plan.prompt.*, plan.rag.*, plan.reviewer_mode.*, plan.retry.*
  • story-ship.mjs: emits ship.* and story_ship.* events
  • refine.mjs: emits refine.analyze, refine.apply, refine.ready
  • guardrails-audit.mjs: emits guardrails.off, guardrails.on, guardrails.auto_shipped
  • auto-phase.mjs: emits auto_phase.transition, auto_phase.invalid

7.6 Telemetry Tools (MCP Query & Reporting)

  • rks_telemetry_query: Filter events by type, date range, correlationId, or runId.
  • rks_telemetry_report: Aggregate summaries, failure breakdowns, and daily trends.

Operational example: A story ship emits a sequence of events:

  1. story_ship.start (with storyId, branch)
  2. story_ship.step.completed (for working_pr)
  3. story_ship.step.completed (for working_merge)
  4. story_ship.step.completed (for mark_implemented)
  5. story_ship.step.completed (for cycle_complete)
  6. story_ship.success (with PR URL, commit IDs)

These events help teams answer questions like: who triggered the action, which validations failed, when did a self-heal attempt run, and which agent produced the plan.

Telemetry Flow

Rendering diagram...

8. Knowledge Management: Dendron + RAG

RKS integrates Dendron notes as a structured knowledge source and uses RAG (retrieval-augmented generation) to provide context to planners and executors.

  • Backlog stories, architecture notes, and policies live in Dendron. Each change to notes can trigger an auto-embed pipeline that updates the semantic index.
  • The planner queries RAG during rks_plan to ensure suggested actions align with constraints and existing patterns.
  • On commit, a rag-embed-on-commit hook refreshes embeddings for changed notes so the knowledge base remains up-to-date.

This combination ensures that plans are not generated in a vacuum: they are grounded in project history, style guides, and prior decisions.

End-to-End Workflow with Embeddings

Rendering diagram...

9. Putting It Together: An End-to-End Example

A typical scenario:

  1. A backlog story is created in Dendron and assigned a problemId.
  2. An agent calls rks_plan(problemId) which queries RAG for context and produces a plan with actionable steps.
  3. rks_exec creates an isolated context and applies changes step-by-step, emitting exec.step events.
  4. After apply, validations run. If all pass, rks_exec commits changes.
  5. rks_story_ship atomically opens a staging PR, merges it, marks the story as integrated, and returns to staging.
  6. MCP gov checks (gov_lint_check, gov_build_check, gov_scope_validate) run on the staging PR.
  7. Release tooling (rks_release) promotes to production according to release policy.

10. Conclusion: Benefits and Tradeoffs

Benefits

  • Predictable, auditable agent-driven operations across any domain
  • Automated validations reduce human review load while preserving quality
  • Traceability through telemetry and correlation IDs
  • Knowledge-grounded plans via Dendron + RAG
  • Multi-layered protection against infrastructure failures (hooks, CI, startup checks)
  • Governed escape hatches for legitimate off-rail work
  • Domain-agnostic patterns applicable to code, business processes, trading, and beyond

Tradeoffs

  • Initial setup and governance policy design require effort
  • Learning curve for teams adopting MCP and hook patterns
  • Some workflows may feel constrained compared to free-form agent operation

When to Use RKS

RKS is a strong fit for teams that need repeatable, auditable automation—whether for software development, business process automation, or any domain where AI agents need guardrails. It's appropriate when you need traceability, human oversight checkpoints, and the ability to audit agent decisions. For one-off prototyping where auditability is not required, lighter weight approaches may be sufficient.

Appendix: Example Telemetry Event (JSON)

{
  "id": "6b8f1a8e-1d5a-4d3b-8f4c-123456789abc",
  "type": "exec.complete",
  "timestamp": "2026-01-22T15:30:00Z",
  "projectId": "routekit-shell",
  "correlationId": "c2d4b6e8-9f01-4a8a-abcdef123456",
  "runId": "rks-run-20260122-001",
  "payload": { "stepsApplied": 5, "latencyMs": 1234 },
  "context": { "branch": "feature/auto-refactor-42", "actor": "agent/auto-123" }
}

This post is meant to be a technical, neutral deep dive for engineering teams evaluating agent orchestration patterns. It focuses on concrete behaviors, function names, and event models to help implementers reason about integrations and tradeoffs. Use the diagrams and governance patterns as templates for adapting RKS principles to your environment—whether that's a software codebase, a trading system, or a document workflow.


Last updated: 2026-02-02 - Broadened framing beyond code-only to include business automation, trading, and other AI agent domains. Added guardrails off/on governance, multi-layered hooks protection, rks_story_ship, CREATE FILE directive support, and plan readiness validation details.