How this works, and how well — sources, licence, measured numbers
What this is, and what it is not
The corpus is 6,397 shell-utility documentation pages — tldr-pages
examples plus a thin slice of Ubuntu 24.04 man pages. Each page was rewritten by an LLM to add a
plain-language summary and the goal-level phrasings someone would type when they want that
tool without knowing its name. That rewrite is the whole trick: documentation is written in the
vocabulary of the tool, and requests arrive in the vocabulary of the goal. fcrackzip's page
says "dictionary attack"; you type "recover the password".
Measured
| corpus / retriever | gold utility in the top 3 |
| original pages, BM25 only | 0.311 |
| rewritten pages, BM25 only | 0.427 |
| rewritten pages, hybrid (this page, 1-bit codes) | 0.439 |
| rewritten pages, hybrid, full-precision server-side | 0.494 |
164 held-out requests whose wording came from a different model than wrote the corpus, over 132
distinct utilities. The lift also survives on 300 requests written by human annotators, so it is not
two language models agreeing with each other. Roughly two in five queries put the right tool in the
top three — useful for finding a utility, not a substitute for reading its actual page.
Credits and licence
Examples are reproduced verbatim from tldr-pages,
© the tldr-pages team and contributors, CC-BY-4.0;
man-page material carries each package's own licence. The generated summaries and phrasings are
unreviewed model output — a retrieval aid, not documentation to trust over the real page.
Index format remax_kb; encoder
MongoDB/mdbr-leaf-mt (int8 ONNX, 23.9 MB);
runtime onnxruntime-web.