dreeseaw/mdlens-realdocs-v1
mdlens-realdocs-v1 A held-out Markdown QA / retrieval eval built entirely from real open-source project documentation. It measures whether an agent can answer documentation questions from the right evidence with fewer irrelevant reads and fewer tokens. Questions are deliberately low lexical overlap (paraphrased), so they stress retrieval rather than string matching. Every non-abstention question has its answer keywords verified to appear in the cited source file.… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/mdlens-realdocs-v1.
0597
