dsergio-hf/boyscouts-handbook-1911-qrels
Boy Scouts Handbook: The First Edition, 1911 A small retrieval and RAG evaluation corpus derived from the Boy Scouts Handbook: The First Edition, 1911, available through Project Gutenberg as eBook #29558. Source Source: Project Gutenberg eBook #29558 Title: Boy Scouts Handbook: The First Edition, 1911 Publisher: Boy Scouts of America Year: 1911 Source URL: https://www.gutenberg.org/ebooks/29558 Source status: Public domain in the United States Corpus… See the full description on the dataset page: https://huggingface.co/datasets/dsergio-hf/boyscouts-handbook-1911-qrels.
Boy Scouts Handbook: The First Edition, 1911
A small retrieval and RAG evaluation corpus derived from the Boy Scouts Handbook: The First Edition, 1911, available through Project Gutenberg as eBook #29558.
Source
- Source: Project Gutenberg eBook #29558
- Title: Boy Scouts Handbook: The First Edition, 1911
- Publisher: Boy Scouts of America
- Year: 1911
- Source URL: https://www.gutenberg.org/ebooks/29558
- Source status: Public domain in the United States
Corpus
The source text is parsed into nine chapter-level documents:
- Scoutcraft
- Woodcraft
- Campcraft
- Tracks, Trailing and Signaling
- Health and Endurance
- Chivalry
- First Aid and Life Saving
- Games and Athletic Standards
- Patriotism and Citizenship
Each document has an identifier, title, and text. Documents are subsequently chunked for retrieval experiments using configurable chunk size and overlap.
Queries and Relevance Judgments
The dataset includes manually constructed retrieval queries and manually assigned relevance judgments (qrels). Relevance is assigned at the chapter/document level rather than to individual retrieval chunks.
This allows different chunking configurations to be evaluated against the same underlying corpus and relevance judgments.
Intended Use
The corpus is intended for:
- semantic retrieval experiments
- retrieval-augmented generation (RAG)
- chunk-size and chunk-overlap experiments
- evaluation of retrieval metrics such as Hit@K and Recall@K
- experiments in grounded question answering
Limitations
This is a historical text from 1911. Its terminology, recommendations, social norms, and safety guidance reflect the period in which it was written and should not be interpreted as contemporary guidance.
In particular, retrieval from this corpus should not be treated as a source of current medical, first-aid, legal, safety, or scouting recommendations.
Provenance
The corpus was constructed from the Project Gutenberg plain-text edition. The source text was parsed into chapter-level documents before being converted to Parquet.
Queries and qrels were created specifically for retrieval experimentation and are not part of the original publication.
