02 - Anatomy of a Retrieval: How Smart Routing Actually Works
Part 1: Technical breakdown of how the Guardrailed Retriever Stack actually works
Anatomy of a Retrieval: How Smart Routing Actually Works
Part 2 of the Technical Deep Dive Series
The Problem in Code
Let's start with the actual problem. Here's what happens when you ask a generic AI about your project:
// Generic AI approach
const response = await ai.complete({
prompt: "How does user authentication work in this project?",
model: "gpt-4"
});
// Result: Generic OAuth2 explanation, no project context
Here's what happens with contextual retrieval:
// Contextual retrieval approach
const context = await retrieveWithRouting(
"How does user authentication work in this project?",
config
);
// Result: Your actual auth implementation, design decisions, and patterns
The Core Router: Decision Engine
The heart of the system is surprisingly simple. Here's the actual routing logic:
// src/router.js - The brain of the operation
export async function retrieveWithRouting(query, cfg, guard) {
// Step 1: Classify the query
const start = classify(query, cfg);
// Step 2: Execute primary search strategy
if (start === "fs") {
fsHits = await fsSearch(query, opts);
// If filesystem search is weak, escalate to semantic search
if (shouldEscalate(fsHits, cfg)) {
ragHits = await ragSearch(query, opts);
}
} else {
// Start with semantic search
ragHits = await ragSearch(query, opts);
// If semantic is weak, try exact match
if (shouldEscalate(ragHits, cfg)) {
fsHits = await fsSearch(query, opts);
}
}
// Step 3: Merge and rank results
let merged = dedupeAndRank([...fsHits, ...ragHits], cfg, guard);
// Step 4: Apply canonical boosting
merged = enforceCanon(merged, cfg);
return merged;
}
The Classification Logic: Smart Query Analysis
Here's how the system decides which search to use first:
function classify(query, cfg) {
// Patterns that suggest filesystem search
const pathPattern = /\.(js|ts|py|md):\d+/; // Looking for file:line
const errorPattern = /error|fail|exception/i; // Error messages
const codeWords = ["function", "class", "import", "TODO"];
// Patterns that suggest semantic search
const conceptWords = ["pattern", "approach", "architecture", "design"];
const isLongQuery = query.split(/\s+/).length >= 6;
// Decision tree
if (pathPattern.test(query)) return "fs";
if (errorPattern.test(query)) return "fs";
if (codeWords.some(w => query.includes(w))) return "fs";
if (conceptWords.some(w => query.includes(w))) return "rag";
if (isLongQuery) return "rag"; // Longer = more semantic
return "fs"; // Default to fast exact search
}
The Filesystem Searcher: Surgical Precision
When you need exact code matches:
// src/retrievers/fs.js
export async function fsSearch(query, opts) {
const searchDirs = ["notes", "src"].filter(dir =>
fs.existsSync(path.resolve(dir))
);
// Use ripgrep for blazing fast search
const args = [
"-n", // Line numbers
"-H", // Filenames
"-C", "2", // 2 lines of context
"--no-ignore",
"--hidden",
query,
...searchDirs
];
const { stdout } = await execa("rg", args, {
timeout: opts.t
});
// Parse ripgrep output into structured results
return parseRipgrepOutput(stdout).map(result => ({
source: "fs",
path: result.path,
text: result.text,
score: 0.5 + contextBonus(result), // Heuristic scoring
}));
}
The Semantic Searcher: Understanding Meaning
When you need conceptual understanding:
// src/retrievers/rag.js
export async function ragSearch(query, opts) {
const { ragDbPath } = getProjectContext();
// Connect to vector database
const db = await connect({
uri: ragDbPath
});
// Generate embedding for query
const queryEmbedding = await embedText(query);
// Find semantically similar documents
const results = await db
.search(queryEmbedding)
.limit(opts.k)
.execute();
return results.map(match => ({
source: "rag",
path: match.path,
text: match.text,
score: match._distance, // Cosine similarity
ts: match.updatedAt
}));
}
The Escalation Logic: Know When to Pivot
The system knows when its first approach isn't working:
function shouldEscalate(results, cfg) {
// No results? Definitely escalate
if (results.length === 0) return true;
// All results have low confidence? Escalate
const avgScore = results.reduce((sum, r) =>
sum + r.score, 0) / results.length;
if (avgScore < cfg.thresholds.escalation_threshold) {
return true;
}
// Results seem good, stick with them
return false;
}
The Canonical Boosting: Trust Authoritative Sources
Some documents are more trustworthy than others:
function enforceCanon(results, cfg) {
return results.map(result => {
// Check if this path matches canonical patterns
const isCanonical = cfg.canonical.paths.some(pattern => {
const regex = new RegExp(pattern.replace("**", ".*"));
return regex.test(result.path);
});
if (isCanonical) {
// Boost canonical sources
result.score += cfg.canonical.boost;
result.canonical = true;
}
return result;
}).sort((a, b) => b.score - a.score);
}
Real Query Walkthrough
Let's trace an actual query through the system:
Query: "How do we handle authentication errors?"
Step 1: Classification
classify("How do we handle authentication errors?")
// Detects "errors" → returns "fs"
Step 2: Filesystem Search
rg -n -H -C 2 "How do we handle authentication errors?" notes src
# Finds 2 matches in error handling code
Step 3: Check Escalation
shouldEscalate(fsResults, cfg)
// 2 results with decent scores → false, no escalation needed
Step 4: Apply Canonical Boost
// decisions.003-auth-error-handling.md gets boosted
// Final ranking puts decision doc first
Result: User gets actual auth error handling code + architectural decision
Query: "What are the best practices for component design?"
Step 1: Classification
classify("What are the best practices for component design?")
// Detects "practices", "design", long query → returns "rag"
Step 2: Semantic Search
// Embedding generated for query
// Searches 233 documentation embeddings
// Finds 8 semantically related documents
Step 3: Check Escalation
shouldEscalate(ragResults, cfg)
// 8 high-quality semantic matches → false
Result: User gets design patterns, component guidelines, best practices
The Configuration: Tuning the System
Here's the actual configuration that controls everything:
# .routekit/retrieval.config.yaml
routing:
fs_triggers:
- regex: "\\.(js|ts|py|md):\\d+" # file:line patterns
- regex: "error|fail|exception" # error patterns
- contains_any: ["function", "class", "import", "TODO"]
rag_triggers:
- contains_any: ["pattern", "approach", "architecture"]
- min_words: 6 # Longer queries are more semantic
budget:
fs_first: { k: 10, time_ms: 5000 }
rag_first: { k: 10, time_ms: 10000 }
thresholds:
escalation_threshold: 0.3 # Escalate if avg score below this
max_total_passages: 20 # Return at most 20 results
canonical:
paths: ["decisions.**", "atoms.**"] # Authoritative sources
boost: 0.2 # Score boost for canonical docs
Performance Characteristics
Here's what actually happens under the hood:
| Search Method | Startup | Search Time | Memory Usage | Accuracy |
|---|---|---|---|---|
| Filesystem Search (ripgrep) | ~10ms | 30-50ms | Minimal (streaming) | 100% for exact matches |
| Semantic Search (LanceDB) | 3-4s (first query only) | 100-200ms | ~500MB (model cached) | 85-95% for semantic similarity |
| Hybrid Routing Performance | Time |
|---|---|
| Classification | <1ms |
| Primary Search | 50-300ms |
| Escalation Decision | <1ms |
| Secondary Search | 50-300ms (if needed) |
| Merge & Rank | <10ms |
| Total Response Time | 100-600ms typical |
The Magic: It's Not Magic
The power isn't in complex algorithms. It's in:
- Smart Classification: Right tool for the right query
- Fast Escalation: Pivot quickly when needed
- Canonical Boosting: Trust authoritative sources
- Simple Scoring: Basic heuristics that work
The entire router is ~200 lines of code. The retrievers are ~100 lines each. The configuration is ~30 lines of YAML.
Simple systems, intelligently combined, create powerful outcomes.
Continue the Series
Next: 03 - Agent Decision Trees - How agents decide what to do with retrieved context
*This is Part 2 of our Technical Deep Dive series exploring AI-first development frameworks. Next:
