hakari-bench/NanoMIRACL
NanoMIRACL This dataset is a Nano-style retrieval dataset for HAKARI-bench. NanoMIRACL is a compact multilingual benchmark derived from MIRACL. Each language split evaluates monolingual retrieval of Wikipedia passages for natural-language questions. This rebuild uses hotchpotch/miracl-hf-unified dev queries at source revision 21ad00eb467639e927b5badb7c49f4947c6c24ca. For each sampled query, it preserves all source positive passages and expands the split-local corpus to 10,000… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/NanoMIRACL.
NanoMIRACL
This dataset is a Nano-style retrieval dataset for HAKARI-bench.
NanoMIRACL is a compact multilingual benchmark derived from MIRACL. Each language split evaluates monolingual retrieval of Wikipedia passages for natural-language questions.
This rebuild uses hotchpotch/miracl-hf-unified dev queries at source revision 21ad00eb467639e927b5badb7c49f4947c6c24ca. For each sampled query, it preserves all source positive passages and expands the split-local corpus to 10,000 documents using source negatives as hard-negative candidates plus deterministic random fill.
Usage
from datasets import load_dataset
dataset_id = "hakari-bench/NanoMIRACL"
split = "ar"
queries = load_dataset(dataset_id, "queries", split=split)
corpus = load_dataset(dataset_id, "corpus", split=split)
qrels = load_dataset(dataset_id, "qrels", split=split)
reranking_candidates = load_dataset(dataset_id, "reranking_hybrid", split=split)Data Layout
This dataset uses six Hugging Face Datasets configs:
corpus: documents with_idandtextqueries: queries with_idandtextqrels: positive relevance labels withquery-idandcorpus-idbm25: BM25 candidate lists withquery-idandcorpus-idsharrier_oss_v1_270m: dense candidate lists frommicrosoft/harrier-oss-v1-270mreranking_hybrid: RRF candidate lists built frombm25andharrier_oss_v1_270m
Each config has the same Nano split names.
Candidate Construction
bm25: local BM25 top-500 with automatic language-aware tokenization. The resolved tokenizer is shown in the Candidate Quality table, for examplewordseg@ja.harrier_oss_v1_270m: dense top-500 frommicrosoft/harrier-oss-v1-270m. In tables this is shown asDense; Dense meansmicrosoft/harrier-oss-v1-270mwith theweb_search_queryprompt for queries and cosine similarity over normalized embeddings.reranking_hybrid: RRF overbm25andharrier_oss_v1_270musingrrf_k=100, keeping the RRF top-100.
Safeguard means rank 101 is appended only when RRF top-100 contains no qrels-positive document.
Split Statistics
Length statistics are character counts computed with len(str(text)).
Candidate Quality
nDCG@10 and Recall@100 are computed from the included candidate rankings against the included qrels, then reported as 0-100 scores such as 52.45. Recall@100 uses only the top 100 candidates; an optional rank-101 safeguard positive is not counted in Recall@100.
Dense means microsoft/harrier-oss-v1-270m with the web_search_query prompt and cosine similarity.
Hybrid Safeguard Summary
- Safeguard positives: 19
- Rows limited by corpus size: 0
- Metadata file:
reranking_hybrid_metadata.json
Source Links
License
NanoMIRACL is a derived dataset. Users must comply with the licenses, terms, and attribution requirements of the upstream source datasets.
