hakari-bench/NanoBEIR-pt
NanoBEIR-pt This dataset is a Nano-style retrieval dataset for HAKARI-bench. NanoBEIR-pt is the Portuguese language-specific component of MNanoBEIR. It groups compact BEIR-derived retrieval tasks for efficient evaluation of document ranking in that language. Usage from datasets import load_dataset dataset_id = "hakari-bench/NanoBEIR-pt" split = "NanoArguAna" queries = load_dataset(dataset_id, "queries", split=split) corpus = load_dataset(dataset_id, "corpus"… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/NanoBEIR-pt.
NanoBEIR-pt
This dataset is a Nano-style retrieval dataset for HAKARI-bench.
NanoBEIR-pt is the Portuguese language-specific component of MNanoBEIR. It groups compact BEIR-derived retrieval tasks for efficient evaluation of document ranking in that language.
Usage
from datasets import load_dataset
dataset_id = "hakari-bench/NanoBEIR-pt"
split = "NanoArguAna"
queries = load_dataset(dataset_id, "queries", split=split)
corpus = load_dataset(dataset_id, "corpus", split=split)
qrels = load_dataset(dataset_id, "qrels", split=split)
reranking_candidates = load_dataset(dataset_id, "reranking_hybrid", split=split)Data Layout
This dataset uses six Hugging Face Datasets configs:
corpus: documents with_idandtextqueries: queries with_idandtextqrels: positive relevance labels withquery-idandcorpus-idbm25: BM25 candidate lists withquery-idandcorpus-idsharrier_oss_v1_270m: dense candidate lists frommicrosoft/harrier-oss-v1-270mreranking_hybrid: RRF candidate lists built frombm25andharrier_oss_v1_270m
Each config has the same Nano split names. NanoNFCorpus includes the full positive qrels (2,518 rows); qrels are not capped to the top-100 reranking depth.
Candidate Construction
bm25: local BM25 top-500 with automatic tokenizer selection. Auto mode useswordsegforja,zh,th,ko, andvi, andregexotherwise. The resolved tokenizer is shown for each split in the Candidate Quality table.harrier_oss_v1_270m: dense top-500 frommicrosoft/harrier-oss-v1-270m. In tables this is shown asDense; Dense meansmicrosoft/harrier-oss-v1-270mwith theweb_search_queryprompt for queries and cosine similarity over normalized embeddings.reranking_hybrid: RRF overbm25andharrier_oss_v1_270musingrrf_k=100, keeping the RRF top-100.
Safeguard means rank 101 is appended only when RRF top-100 contains no qrels-positive document. Qrels are not capped to fit the top-100 reranking depth. For NanoNFCorpus, some queries have more than 100 positive qrels, so top-100 hybrid candidate coverage is expected to be below 100%; this is a candidate-list diagnostic, not a qrels filtering rule.
Split Statistics
Length statistics are character counts computed with len(str(text)).
Candidate Quality
nDCG@10 and Recall@100 are computed from the included candidate rankings against the included qrels, then reported as 0-100 scores such as 52.45. Recall@100 uses only the top 100 candidates; an optional rank-101 safeguard positive is not counted in Recall@100.
Dense means microsoft/harrier-oss-v1-270m with the web_search_query prompt and cosine similarity.
Hybrid Safeguard Summary
- Safeguard positives: 26
- Rows limited by corpus size: 0
- Metadata file:
reranking_hybrid_metadata.json
Source Links
- Original dataset: lightonai/NanoBEIR-pt
- Final dataset: hakari-bench/NanoBEIR-pt
License
NanoBEIR-pt is a derived dataset. Users must comply with the licenses, terms, and attribution requirements of the upstream source datasets.
