datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scidocs_colbert_pgtr_golden
SciDocs ColBERT PGTR Golden
SciDocs queries and corpus from Nithish2410/scidocs_colbert_pgtr_golden, with the previous targets ignored and replaced by full Qwen-reranked top-100 targets.
Contents
train.jsonl: 14,142 queries with 100 Qwen-reranked targets each.
items.jsonl: 25,657 SciDocs corpus passages.
Rerank Setup
Query source: existing query texts from the dataset.
Corpus source: existing items split from the dataset.
Candidate source:… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/scidocs_colbert_pgtr_golden.arxiv_colbert_pgtr_golden
ArXiv ColBERT PGTR Golden
ArXiv queries and corpus from Nithish2410/arxiv_colbert_pgtr_golden, with the previous targets ignored and replaced by Qwen-reranked top-100 targets.
Contents
train.jsonl: 10,000 queries with 100 Qwen-reranked targets each.
items.jsonl: 2,040 ArXiv corpus passages.
Rerank Setup
Query source: existing query texts from the dataset.
Corpus source: existing items split from the dataset.
Candidate source: e5-base-v2 retrieval… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/arxiv_colbert_pgtr_golden.covid_colbert_pgtr_golden
COVID ColBERT PGTR Golden
COVID queries and corpus from Nithish2410/covid_colbert_pgtr_golden, with the previous targets ignored and replaced by Qwen-reranked top-100 targets.
Contents
train.jsonl: 10,000 queries with 100 Qwen-reranked targets each.
items.jsonl: 171,332 COVID corpus passages.
Rerank Setup
Query source: existing query texts from the dataset.
Corpus source: existing items split from the dataset.
Candidate source: e5-base-v2… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/covid_colbert_pgtr_golden.multi-lang-reranker-colbertcolbert-wiki17
