CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PeterJinGo /wiki-18-bm25-index0 likes5.2k downloads1y agoHugging Face02sentence-transformers /msmarco-bm25 MS MARCO with hard negatives from bm25 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-bm25.tabularfeature-extraction10M<n<100M4 likes2.9k downloads2y agoHugging Face03princeton-nlp /SWE-bench_bm25_13K Dataset Card for "SWE-bench_bm25_13K" Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? This dataset SWE-bench_bm25_13K includes a formatting of… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_13K.text10K<n<100K3 likes1.3k downloads2y agoHugging Face04princeton-nlp /SWE-bench_bm25_40K Dataset Card for "SWE-bench_bm25_40K" Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? This dataset SWE-bench_bm25_40K includes a formatting of… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_40K.text10K<n<100K3 likes1.3k downloads2y agoHugging Face05princeton-nlp /SWE-bench_bm25_27K Dataset Card for "SWE-bench_bm25_27K" Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? This dataset SWE-bench_bm25_27K includes a formatting of… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_27K.text10K<n<100K1 likes1.3k downloads2y agoHugging Face06eth-sri /SWT-bench_Lite_bm25_27k_zsb Dataset Summary SWT-bench Lite is subset of SWT-bench, a dataset that tests systems’ ability to reproduce GitHub issues automatically. The dataset collects 276 test Issue-Pull Request pairs from 11 popular Python GitHub projects. Evaluation is performed by unit test verification using pre- and post-PR behavior of the test suite with and without the model proposed tests. 📊🏆 Leaderboard A public leaderboard for performance on SWT-bench is hosted at swtbench.com The… See the full description on the dataset page: https://huggingface.co/datasets/eth-sri/SWT-bench_Lite_bm25_27k_zsb.textn<1K0 likes445 downloads2y agoHugging Face07princeton-nlp /SWE-bench_bm25_50k_llama Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. Supported Tasks and Leaderboards SWE-bench proposes a new task: issue resolution provided a full repository and GitHub issue. The leaderboard can be found at www.swebench.com Languages… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_50k_llama.text1K<n<10K6 likes418 downloads2y agoHugging Face08Maktabati /openiti-bm25 OpenITI BM25 Index BM25 full-text search index for the OpenITI corpus (release 2025.1.9), built on SQLite FTS5. Part of the maktabati.ai Islamic RAG pipeline. Designed to pair with Maktabati/openiti-vectors for hybrid retrieval (BM25 + Dense + RRF fusion). Statistics Source OpenITI 2025.1.9 Files indexed ~8,943 (one edition per work) Languages Arabic (ar), Persian (fa), Turkish (tr), Urdu (ur) Chunk size 512 tokens, 50-token overlap Tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/openiti-bm25.1M<n<10M0 likes409 downloads4mo agoHugging Face09princeton-nlp /SWE-bench_Lite_bm25_13K Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? This dataset SWE-bench_Lite_bm25_13K includes a formatting of each instance… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite_bm25_13K.textn<1K1 likes377 downloads2y agoHugging Face10princeton-nlp /SWE-bench_bm25_27k_cl100k Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. Supported Tasks and Leaderboards SWE-bench proposes a new task: issue resolution provided a full repository and GitHub issue. The leaderboard can be found at www.swebench.com Languages… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_27k_cl100k.text1K<n<10K0 likes369 downloads2y agoHugging Face11princeton-nlp /SWE-bench_Lite_bm25_27K Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? This dataset SWE-bench_Lite_bm25_27K includes a formatting of each instance… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite_bm25_27K.textn<1K1 likes350 downloads2y agoHugging Face12princeton-nlp /SWE-bench_bm25_13k_cl100k Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. Supported Tasks and Leaderboards SWE-bench proposes a new task: issue resolution provided a full repository and GitHub issue. The leaderboard can be found at www.swebench.com Languages… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_13k_cl100k.text1K<n<10K0 likes328 downloads2y agoHugging Face13CCDS-Scaling /bbh-train-p1.0-bm25text100K<n<1M0 likes274 downloads2y agoHugging Face14anime-sh /generations-olmo-3-7b-rmu-bm25-6t-rebuttaltabular10K<n<100K0 likes269 downloads2mo agoHugging Face15Comet /wikipedia-2017-bm25 Wikipedia 2017 BM25 Search Index This dataset provides a production-ready BM25 search index over 5.2 million Wikipedia article abstracts from the 2017 snapshot. Built using the bm25s library with English stemming and optimized Parquet compression, it enables fast, offline information retrieval for research and production AI systems. The corpus is identical to the one used in influential AI research papers including DSPy and GEPA, ensuring reproducible benchmarking and fair… See the full description on the dataset page: https://huggingface.co/datasets/Comet/wikipedia-2017-bm25.texttext-retrieval1M<n<10M1 likes263 downloads10mo agoHugging Face16anime-sh /generations-llama-3_1-8b-rmu-bm25-10b-rebuttaltabular10K<n<100K0 likes252 downloads2mo agoHugging Face17Maktabati /shamela-bm25 Shamela BM25 Index BM25 full-text search index for the Al-Maktaba Al-Shamela digital library, built on SQLite FTS5. Part of the maktabati.ai Islamic RAG pipeline. Designed to pair with Maktabati/shamela-vectors for hybrid retrieval (BM25 + Dense + RRF fusion). Statistics Books 8,589 Categories 40 Book chunks 12,331,995 Quran verses (standalone) 6,236 Total rows 12,338,231 Chunking: 512 tokens, 50-token overlap (tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/shamela-bm25.textn<1K0 likes246 downloads4mo agoHugging Face18anime-sh /generations-olmo-3-7b-undial-bm25-6t-rebuttaltabular10K<n<100K0 likes231 downloads2mo agoHugging Face19anime-sh /generations-olmo-3-7b-rmu-bm25-10b-rebuttaltabular10K<n<100K0 likes229 downloads2mo agoHugging Face20Doohae /klue-mrc-bm25text10K<n<100K0 likes226 downloads5y agoHugging Face21anime-sh /generations-llama-3_1-8b-undial-bm25-6t-rebuttaltabular10K<n<100K0 likes222 downloads2mo agoHugging Face22hotchpotch /NanoBEIR-es-with-bm25🚧 This dataset is currently under construction. Specifications may change. NanoBEIR-es (with bm25 subset) A BM25-first retrieval subset derived from lightonai/NanoBEIR-es for reproducible NanoBEIR evaluation. What this dataset is Language: es BM25 top-k: 100 Number of evaluation splits: 13 This folder contains generated BM25 retrieval results and reproducibility metadata. Data structure corpus: _id, text queries: _id, text qrels: query-id, corpus-id, score… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/NanoBEIR-es-with-bm25.text10K<n<100K0 likes220 downloads7mo agoHugging Face23anime-sh /generations-qwen3-8b-undial-bm25-6t-rebuttaltabular10K<n<100K0 likes216 downloads2mo agoHugging Face24anime-sh /generations-llama-3_1-8b-rmu-bm25-6t-rebuttaltabular10K<n<100K0 likes208 downloads2mo agoHugging Face25anime-sh /generations-qwen3-8b-rmu-bm25-6t-rebuttaltabular10K<n<100K0 likes205 downloads2mo agoHugging Face26anime-sh /generations-qwen3-8b-rmu-bm25-10b-rebuttaltabular10K<n<100K0 likes204 downloads2mo agoHugging Face27anime-sh /generations-qwen3-8b-undial-bm25-10b-rebuttaltabular10K<n<100K0 likes203 downloads2mo agoHugging Face28anime-sh /generations-olmo-3-7b-undial-bm25-10b-rebuttaltabular10K<n<100K0 likes195 downloads2mo agoHugging Face29Nithish2410 /benchmark-miriad-200k-bm25-100-q8brerank MIRIAD Benchmark 200k BM25 Top-100 Qwen Rerank MIRIAD 1k benchmark variant using the same 1,000 official test queries and the same 200,000-passage corpus as Nithish2410/benchmark-miriad-200k, but replacing the original single-positive qrels with Qwen reranker scores over BM25 top-100 candidates. Contents test.jsonl: 1,000 test queries with 100 reranked targets each. items.jsonl: 200,000 MIRIAD corpus passages. Label Source Candidate source: BM25… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/benchmark-miriad-200k-bm25-100-q8brerank.text100K<n<1M0 likes193 downloads23d agoHugging Face30hotchpotch /NanoBEIR-th-with-bm25🚧 This dataset is currently under construction. Specifications may change. NanoBEIR-th (with bm25 subset) A BM25-first retrieval subset derived from sionic-ai/NanoBEIR-th for reproducible NanoBEIR evaluation. What this dataset is Language: th BM25 top-k: 100 Number of evaluation splits: 13 This folder contains generated BM25 retrieval results and reproducibility metadata. Data structure corpus: _id, text queries: _id, text qrels: query-id, corpus-id, score… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/NanoBEIR-th-with-bm25.text10K<n<100K0 likes192 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.