
Reproducibility Is a Feature
Corpus, question sets and harness are committed, and every table regenerates with one command. When the headline is 0/11 versus 11/11 for the same model, the only honest answer to suspicion is a rerun.
4 min read
Blog
Engineering notes on navigation, ingestion, agent memory and honest benchmarking, written by the people who ran the experiments.

Corpus, question sets and harness are committed, and every table regenerates with one command. When the headline is 0/11 versus 11/11 for the same model, the only honest answer to suspicion is a rerun.
4 min read

Trail learning missed its convergence criterion: hops fell by roughly half the threshold. The post-mortem: why sharp entry search and disciplined curation left trails almost nothing to compress.
4 min read

Eleven questions, each needing at least three chained hops. The same 12B model scores 0/11 as a top-k RAG reader and 11/11 as a forest navigator. Here is how the benchmark was built and how to rerun it.
4 min read

The trail-learning convergence criterion was not met, and it is in the report anyway. A benchmark you can only pass is not a benchmark.
3 min read
The paper carries the full architecture, the benchmark tables and the findings that failed their criteria.