← Back to Thinking

02 - Anatomy of a Retrieval: How Smart Routing Actually Works

Vince Mease• Technical Deep Dive• 9/6/2025
project:routekit-shellblogtechnicalretrievalimplementation

Part 1: Technical breakdown of how the Guardrailed Retriever Stack actually works

Anatomy of a Retrieval: How Smart Routing Actually Works

Part 2 of the Technical Deep Dive Series

The Problem in Code

Let's start with the actual problem. Here's what happens when you ask a generic AI about your project:

// Generic AI approach
const response = await ai.complete({
  prompt: "How does user authentication work in this project?",
  model: "gpt-4"
});
// Result: Generic OAuth2 explanation, no project context

Here's what happens with contextual retrieval:

// Contextual retrieval approach
const context = await retrieveWithRouting(
  "How does user authentication work in this project?",
  config
);
// Result: Your actual auth implementation, design decisions, and patterns

The Core Router: Decision Engine

The heart of the system is surprisingly simple. Here's the actual routing logic:

// src/router.js - The brain of the operation
export async function retrieveWithRouting(query, cfg, guard) {
  // Step 1: Classify the query
  const start = classify(query, cfg);
  
  // Step 2: Execute primary search strategy
  if (start === "fs") {
    fsHits = await fsSearch(query, opts);
    // If filesystem search is weak, escalate to semantic search
    if (shouldEscalate(fsHits, cfg)) {
      ragHits = await ragSearch(query, opts);
    }
  } else {
    // Start with semantic search
    ragHits = await ragSearch(query, opts);
    // If semantic is weak, try exact match
    if (shouldEscalate(ragHits, cfg)) {
      fsHits = await fsSearch(query, opts);
    }
  }
  
  // Step 3: Merge and rank results
  let merged = dedupeAndRank([...fsHits, ...ragHits], cfg, guard);
  
  // Step 4: Apply canonical boosting
  merged = enforceCanon(merged, cfg);
  
  return merged;
}

The Classification Logic: Smart Query Analysis

Here's how the system decides which search to use first:

function classify(query, cfg) {
  // Patterns that suggest filesystem search
  const pathPattern = /\.(js|ts|py|md):\d+/;  // Looking for file:line
  const errorPattern = /error|fail|exception/i; // Error messages
  const codeWords = ["function", "class", "import", "TODO"];
  
  // Patterns that suggest semantic search  
  const conceptWords = ["pattern", "approach", "architecture", "design"];
  const isLongQuery = query.split(/\s+/).length >= 6;
  
  // Decision tree
  if (pathPattern.test(query)) return "fs";
  if (errorPattern.test(query)) return "fs"; 
  if (codeWords.some(w => query.includes(w))) return "fs";
  
  if (conceptWords.some(w => query.includes(w))) return "rag";
  if (isLongQuery) return "rag"; // Longer = more semantic
  
  return "fs"; // Default to fast exact search
}

The Filesystem Searcher: Surgical Precision

When you need exact code matches:

// src/retrievers/fs.js
export async function fsSearch(query, opts) {
  const searchDirs = ["notes", "src"].filter(dir => 
    fs.existsSync(path.resolve(dir))
  );
  
  // Use ripgrep for blazing fast search
  const args = [
    "-n",        // Line numbers
    "-H",        // Filenames
    "-C", "2",   // 2 lines of context
    "--no-ignore",
    "--hidden",
    query,
    ...searchDirs
  ];
  
  const { stdout } = await execa("rg", args, { 
    timeout: opts.t 
  });
  
  // Parse ripgrep output into structured results
  return parseRipgrepOutput(stdout).map(result => ({
    source: "fs",
    path: result.path,
    text: result.text,
    score: 0.5 + contextBonus(result), // Heuristic scoring
  }));
}

The Semantic Searcher: Understanding Meaning

When you need conceptual understanding:

// src/retrievers/rag.js
export async function ragSearch(query, opts) {
  const { ragDbPath } = getProjectContext();
  
  // Connect to vector database
  const db = await connect({
    uri: ragDbPath
  });
  
  // Generate embedding for query
  const queryEmbedding = await embedText(query);
  
  // Find semantically similar documents
  const results = await db
    .search(queryEmbedding)
    .limit(opts.k)
    .execute();
  
  return results.map(match => ({
    source: "rag",
    path: match.path,
    text: match.text,
    score: match._distance, // Cosine similarity
    ts: match.updatedAt
  }));
}

The Escalation Logic: Know When to Pivot

The system knows when its first approach isn't working:

function shouldEscalate(results, cfg) {
  // No results? Definitely escalate
  if (results.length === 0) return true;
  
  // All results have low confidence? Escalate
  const avgScore = results.reduce((sum, r) => 
    sum + r.score, 0) / results.length;
  
  if (avgScore < cfg.thresholds.escalation_threshold) {
    return true;
  }
  
  // Results seem good, stick with them
  return false;
}

The Canonical Boosting: Trust Authoritative Sources

Some documents are more trustworthy than others:

function enforceCanon(results, cfg) {
  return results.map(result => {
    // Check if this path matches canonical patterns
    const isCanonical = cfg.canonical.paths.some(pattern => {
      const regex = new RegExp(pattern.replace("**", ".*"));
      return regex.test(result.path);
    });
    
    if (isCanonical) {
      // Boost canonical sources
      result.score += cfg.canonical.boost;
      result.canonical = true;
    }
    
    return result;
  }).sort((a, b) => b.score - a.score);
}

Real Query Walkthrough

Let's trace an actual query through the system:

Query: "How do we handle authentication errors?"

Step 1: Classification

classify("How do we handle authentication errors?")
// Detects "errors" → returns "fs"

Step 2: Filesystem Search

rg -n -H -C 2 "How do we handle authentication errors?" notes src
# Finds 2 matches in error handling code

Step 3: Check Escalation

shouldEscalate(fsResults, cfg)
// 2 results with decent scores → false, no escalation needed

Step 4: Apply Canonical Boost

// decisions.003-auth-error-handling.md gets boosted
// Final ranking puts decision doc first

Result: User gets actual auth error handling code + architectural decision

Query: "What are the best practices for component design?"

Step 1: Classification

classify("What are the best practices for component design?")
// Detects "practices", "design", long query → returns "rag"

Step 2: Semantic Search

// Embedding generated for query
// Searches 233 documentation embeddings
// Finds 8 semantically related documents

Step 3: Check Escalation

shouldEscalate(ragResults, cfg)
// 8 high-quality semantic matches → false

Result: User gets design patterns, component guidelines, best practices

The Configuration: Tuning the System

Here's the actual configuration that controls everything:

# .routekit/retrieval.config.yaml
routing:
  fs_triggers:
    - regex: "\\.(js|ts|py|md):\\d+"  # file:line patterns
    - regex: "error|fail|exception"    # error patterns
    - contains_any: ["function", "class", "import", "TODO"]
  
  rag_triggers:
    - contains_any: ["pattern", "approach", "architecture"]
    - min_words: 6  # Longer queries are more semantic

budget:
  fs_first: { k: 10, time_ms: 5000 }
  rag_first: { k: 10, time_ms: 10000 }

thresholds:
  escalation_threshold: 0.3  # Escalate if avg score below this
  max_total_passages: 20      # Return at most 20 results

canonical:
  paths: ["decisions.**", "atoms.**"]  # Authoritative sources
  boost: 0.2  # Score boost for canonical docs

Performance Characteristics

Here's what actually happens under the hood:

Search Method Startup Search Time Memory Usage Accuracy
Filesystem Search (ripgrep) ~10ms 30-50ms Minimal (streaming) 100% for exact matches
Semantic Search (LanceDB) 3-4s (first query only) 100-200ms ~500MB (model cached) 85-95% for semantic similarity

Hybrid Routing Performance Time
Classification <1ms
Primary Search 50-300ms
Escalation Decision <1ms
Secondary Search 50-300ms (if needed)
Merge & Rank <10ms
Total Response Time 100-600ms typical

The Magic: It's Not Magic

The power isn't in complex algorithms. It's in:

  1. Smart Classification: Right tool for the right query
  2. Fast Escalation: Pivot quickly when needed
  3. Canonical Boosting: Trust authoritative sources
  4. Simple Scoring: Basic heuristics that work

The entire router is ~200 lines of code. The retrievers are ~100 lines each. The configuration is ~30 lines of YAML.

Simple systems, intelligently combined, create powerful outcomes.

Continue the Series

Next: 03 - Agent Decision Trees - How agents decide what to do with retrieved context


*This is Part 2 of our Technical Deep Dive series exploring AI-first development frameworks. Next: