Skip to content

Empirical Benchmarks & Evaluation

Waymark Engine has been rigorously evaluated against standard agentic retrieval methods, including whole-repo vector embeddings and Aider's AST PageRank repo map.


Steelmanned Aider Repo Map Parity Benchmark

Evaluated using gpt-oss-120b across 6 distinct retrieval categories on real-world repositories:

Retrieval Task Aider Repo Map (PageRank) Waymark Engine (4-Tier) Variance
Exact Symbol Caller 1,840 tokens (ranked file dump) 34 tokens (Tier 1 AST hit) -98.1% tokens
Multi-Hop BFS Trace Not supported (requires agent loop) 112 tokens (Tier 1 Bounded BFS) Deterministic graph
Typo / Transposition 0% recall (miss or wrong file) 100% recall (Tier 3 Junegunn FZF) +100% resilience
Conceptual Invariant 2,400 tokens (approximate text) 86 tokens (Tier 4 BM25 Consensus) -96.4% tokens
False-Positive Defense Hallucinates related functions 0 false positives (fail-closed miss) Zero hallucinations

Token Economics & Efficiency

When AI coding agents retrieve context, large file dumps rapidly degrade reasoning quality and inflate token costs:

Full File Context Reads (Standard Agent):
████████████████████████████████████████  15,000–45,000 tokens

Aider Repo Map Ranking:
████████████  2,500–5,000 tokens

Waymark Engine (--plain):
█  16–38 tokens (99.8% reduction)

Test Exclusion Impact

Call graphs in large repositories are often overwhelmed by unit and integration tests. Enabling --exclude-tests: - Suppresses test helper fixtures, mock callers, and test files (*_test.*, *.spec.*, tests/). - Yields up to 94% token savings on multi-hop caller trees while preserving 100% of production architectural paths.


Latency & Performance Benchmarks

Measured on an AMD Ryzen 9 7950X / Windows 11 with Node.js v22 across a 50,000-line TypeScript codebase:

Operation Cold Process Scan Warm Resident Daemon Speedup
Tier 1 (AST Symbol Definition) 185 ms 8.2 ms 22.5x
Tier 1 (Bounded Multi-Hop BFS) 240 ms 12.4 ms 19.3x
Tier 2 (Literal Path PrefixTrie) 24 ms 0.8 ms 30.0x
Tier 3 (Junegunn Choi FZF Score) 42 ms 3.1 ms 13.5x
Tier 4 (Lexical BM25 Memory Hit) 35 ms 5.4 ms 6.5x

Adversarial Evaluation Battery

Waymark maintains an automated 6-category regression and adversarial evaluation battery (docs/prompts/ and /battery in the REPL):

  1. AST Call Graph Fidelity: Deterministic verification of callers and callees across overloaded signatures.
  2. Literal Path Resolution: Absolute, relative, and ambiguous path disambiguation.
  3. Fuzzy Scoring Accuracy: Transposition resilience, camelCase boundary weighting, and separator biases.
  4. Bounded Multi-Hop Invariants: Cycle detection, depth limiting (1–5), and directionality enforcement.
  5. Batch Symbol Concurrency: Concurrent multi-symbol queries under thread bounds.
  6. Fail-Closed Adversarial Resistance: Zero hallucinations on fabricated symbols, non-existent files, and uninitialized vector stores.

All tests execute with 100% pass rates across offline test suites and live CI runs.