datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wiki-18-bm25-indexmsmarco-bm25
MS MARCO with hard negatives from bm25
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:
msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-bm25.SWE-bench_bm25_13K
Dataset Card for "SWE-bench_bm25_13K"
Dataset Summary
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
This dataset SWE-bench_bm25_13K includes a formatting of… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_13K.SWE-bench_bm25_40K
Dataset Card for "SWE-bench_bm25_40K"
Dataset Summary
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
This dataset SWE-bench_bm25_40K includes a formatting of… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_40K.SWE-bench_bm25_27K
Dataset Card for "SWE-bench_bm25_27K"
Dataset Summary
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
This dataset SWE-bench_bm25_27K includes a formatting of… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_27K.SWT-bench_Lite_bm25_27k_zsb
Dataset Summary
SWT-bench Lite is subset of SWT-bench, a dataset that tests systems’ ability to reproduce GitHub issues automatically. The dataset collects 276 test Issue-Pull Request pairs from 11 popular Python GitHub projects. Evaluation is performed by unit test verification using pre- and post-PR behavior of the test suite with and without the model proposed tests.
📊🏆 Leaderboard
A public leaderboard for performance on SWT-bench is hosted at swtbench.com
The… See the full description on the dataset page: https://huggingface.co/datasets/eth-sri/SWT-bench_Lite_bm25_27k_zsb.SWE-bench_bm25_50k_llama
Dataset Summary
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
Supported Tasks and Leaderboards
SWE-bench proposes a new task: issue resolution provided a full repository and GitHub issue. The leaderboard can be found at www.swebench.com
Languages… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_50k_llama.openiti-bm25
OpenITI BM25 Index
BM25 full-text search index for the OpenITI corpus (release 2025.1.9), built on SQLite FTS5. Part of the maktabati.ai Islamic RAG pipeline.
Designed to pair with Maktabati/openiti-vectors for hybrid retrieval (BM25 + Dense + RRF fusion).
Statistics
Source
OpenITI 2025.1.9
Files indexed
~8,943 (one edition per work)
Languages
Arabic (ar), Persian (fa), Turkish (tr), Urdu (ur)
Chunk size
512 tokens, 50-token overlap
Tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/openiti-bm25.SWE-bench_Lite_bm25_13K
Dataset Summary
SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
This dataset SWE-bench_Lite_bm25_13K includes a formatting of each instance… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite_bm25_13K.SWE-bench_bm25_27k_cl100k
Dataset Summary
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
Supported Tasks and Leaderboards
SWE-bench proposes a new task: issue resolution provided a full repository and GitHub issue. The leaderboard can be found at www.swebench.com
Languages… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_27k_cl100k.SWE-bench_Lite_bm25_27K
Dataset Summary
SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
This dataset SWE-bench_Lite_bm25_27K includes a formatting of each instance… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite_bm25_27K.SWE-bench_bm25_13k_cl100k
Dataset Summary
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
Supported Tasks and Leaderboards
SWE-bench proposes a new task: issue resolution provided a full repository and GitHub issue. The leaderboard can be found at www.swebench.com
Languages… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_13k_cl100k.bbh-train-p1.0-bm25generations-olmo-3-7b-rmu-bm25-6t-rebuttalwikipedia-2017-bm25
Wikipedia 2017 BM25 Search Index
This dataset provides a production-ready BM25 search index over 5.2 million Wikipedia article abstracts from the 2017 snapshot. Built using the bm25s library with English stemming and optimized Parquet compression, it enables fast, offline information retrieval for research and production AI systems. The corpus is identical to the one used in influential AI research papers including DSPy and GEPA, ensuring reproducible benchmarking and fair… See the full description on the dataset page: https://huggingface.co/datasets/Comet/wikipedia-2017-bm25.generations-llama-3_1-8b-rmu-bm25-10b-rebuttalshamela-bm25
Shamela BM25 Index
BM25 full-text search index for the Al-Maktaba Al-Shamela digital library, built on SQLite FTS5. Part of the maktabati.ai Islamic RAG pipeline.
Designed to pair with Maktabati/shamela-vectors for hybrid retrieval (BM25 + Dense + RRF fusion).
Statistics
Books
8,589
Categories
40
Book chunks
12,331,995
Quran verses (standalone)
6,236
Total rows
12,338,231
Chunking: 512 tokens, 50-token overlap (tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/shamela-bm25.generations-olmo-3-7b-undial-bm25-6t-rebuttalgenerations-olmo-3-7b-rmu-bm25-10b-rebuttalklue-mrc-bm25generations-llama-3_1-8b-undial-bm25-6t-rebuttalNanoBEIR-es-with-bm25🚧 This dataset is currently under construction. Specifications may change.
NanoBEIR-es (with bm25 subset)
A BM25-first retrieval subset derived from lightonai/NanoBEIR-es for reproducible NanoBEIR evaluation.
What this dataset is
Language: es
BM25 top-k: 100
Number of evaluation splits: 13
This folder contains generated BM25 retrieval results and reproducibility metadata.
Data structure
corpus: _id, text
queries: _id, text
qrels: query-id, corpus-id, score… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/NanoBEIR-es-with-bm25.generations-qwen3-8b-undial-bm25-6t-rebuttalgenerations-llama-3_1-8b-rmu-bm25-6t-rebuttalgenerations-qwen3-8b-rmu-bm25-6t-rebuttalgenerations-qwen3-8b-rmu-bm25-10b-rebuttalgenerations-qwen3-8b-undial-bm25-10b-rebuttalgenerations-olmo-3-7b-undial-bm25-10b-rebuttalbenchmark-miriad-200k-bm25-100-q8brerank
MIRIAD Benchmark 200k BM25 Top-100 Qwen Rerank
MIRIAD 1k benchmark variant using the same 1,000 official test queries and the same 200,000-passage corpus as Nithish2410/benchmark-miriad-200k, but replacing the original single-positive qrels with Qwen reranker scores over BM25 top-100 candidates.
Contents
test.jsonl: 1,000 test queries with 100 reranked targets each.
items.jsonl: 200,000 MIRIAD corpus passages.
Label Source
Candidate source: BM25… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/benchmark-miriad-200k-bm25-100-q8brerank.NanoBEIR-th-with-bm25🚧 This dataset is currently under construction. Specifications may change.
NanoBEIR-th (with bm25 subset)
A BM25-first retrieval subset derived from sionic-ai/NanoBEIR-th for reproducible NanoBEIR evaluation.
What this dataset is
Language: th
BM25 top-k: 100
Number of evaluation splits: 13
This folder contains generated BM25 retrieval results and reproducibility metadata.
Data structure
corpus: _id, text
queries: _id, text
qrels: query-id, corpus-id, score… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/NanoBEIR-th-with-bm25.
