Empirical Benchmarks & Evaluation¶
Waymark Engine has been rigorously evaluated against standard agentic retrieval methods, including whole-repo vector embeddings and Aider's AST PageRank repo map.
Steelmanned Aider Repo Map Parity Benchmark¶
Evaluated using gpt-oss-120b across 6 distinct retrieval categories on real-world repositories:
| Retrieval Task | Aider Repo Map (PageRank) | Waymark Engine (4-Tier) | Variance |
|---|---|---|---|
| Exact Symbol Caller | 1,840 tokens (ranked file dump) | 34 tokens (Tier 1 AST hit) | -98.1% tokens |
| Multi-Hop BFS Trace | Not supported (requires agent loop) | 112 tokens (Tier 1 Bounded BFS) | Deterministic graph |
| Typo / Transposition | 0% recall (miss or wrong file) | 100% recall (Tier 3 Junegunn FZF) | +100% resilience |
| Conceptual Invariant | 2,400 tokens (approximate text) | 86 tokens (Tier 4 BM25 Consensus) | -96.4% tokens |
| False-Positive Defense | Hallucinates related functions | 0 false positives (fail-closed miss) | Zero hallucinations |
Token Economics & Efficiency¶
When AI coding agents retrieve context, large file dumps rapidly degrade reasoning quality and inflate token costs:
Full File Context Reads (Standard Agent):
████████████████████████████████████████ 15,000–45,000 tokens
Aider Repo Map Ranking:
████████████ 2,500–5,000 tokens
Waymark Engine (--plain):
█ 16–38 tokens (99.8% reduction)
Test Exclusion Impact¶
Call graphs in large repositories are often overwhelmed by unit and integration tests. Enabling --exclude-tests:
- Suppresses test helper fixtures, mock callers, and test files (*_test.*, *.spec.*, tests/).
- Yields up to 94% token savings on multi-hop caller trees while preserving 100% of production architectural paths.
Latency & Performance Benchmarks¶
Measured on an AMD Ryzen 9 7950X / Windows 11 with Node.js v22 across a 50,000-line TypeScript codebase:
| Operation | Cold Process Scan | Warm Resident Daemon | Speedup |
|---|---|---|---|
| Tier 1 (AST Symbol Definition) | 185 ms | 8.2 ms | 22.5x |
| Tier 1 (Bounded Multi-Hop BFS) | 240 ms | 12.4 ms | 19.3x |
| Tier 2 (Literal Path PrefixTrie) | 24 ms | 0.8 ms | 30.0x |
| Tier 3 (Junegunn Choi FZF Score) | 42 ms | 3.1 ms | 13.5x |
| Tier 4 (Lexical BM25 Memory Hit) | 35 ms | 5.4 ms | 6.5x |
Adversarial Evaluation Battery¶
Waymark maintains an automated 6-category regression and adversarial evaluation battery (docs/prompts/ and /battery in the REPL):
- AST Call Graph Fidelity: Deterministic verification of callers and callees across overloaded signatures.
- Literal Path Resolution: Absolute, relative, and ambiguous path disambiguation.
- Fuzzy Scoring Accuracy: Transposition resilience, camelCase boundary weighting, and separator biases.
- Bounded Multi-Hop Invariants: Cycle detection, depth limiting (1–5), and directionality enforcement.
- Batch Symbol Concurrency: Concurrent multi-symbol queries under thread bounds.
- Fail-Closed Adversarial Resistance: Zero hallucinations on fabricated symbols, non-existent files, and uninitialized vector stores.
All tests execute with 100% pass rates across offline test suites and live CI runs.