
The Usual Patches, and Where They Stop
Rerankers, rewriting, HyDE, loops and agents all help, which is why they are everywhere. Here is what each one costs, where each one stops on multi-hop, and the gap that remains, measured.
4 min read
Blog
Engineering notes on navigation, ingestion, agent memory and honest benchmarking, written by the people who ran the experiments.

Rerankers, rewriting, HyDE, loops and agents all help, which is why they are everywhere. Here is what each one costs, where each one stops on multi-hop, and the gap that remains, measured.
4 min read

Per-query cost makes failure look like a discount. Measured per correct answer, the forest navigator spends 0.58x the tokens of an iterative-RAG baseline, same 12B model. Here is the arithmetic.
4 min read

Eleven questions, each needing at least three chained hops. The same 12B model scores 0/11 as a top-k RAG reader and 11/11 as a forest navigator. Here is how the benchmark was built and how to rerun it.
4 min read

One question, three documents, traced twice: watch top-k retrieval dead-end on a three-hop question, then watch the same corpus answer it when an agent walks it node by node.
4 min read

Three moves, one worked hunt through a small company's corpus, and why the same 12B model goes from 0/11 to 11/11 when it walks a forest instead of reading a top-k paste.
4 min read

Iterative RAG loops hide their cost in the wrong denominator. Measured per correct answer: 0.58x the tokens and 8.4 s p95 vs 17.5 s, same 12B model.
3 min read

Top-k retrieval is a single hop by construction. When an answer needs three, the bottleneck stops being the model and becomes the shape of your corpus.
4 min read
The paper carries the full architecture, the benchmark tables and the findings that failed their criteria.