bm25
Datasets
All datasets matching “bm25”wiki-18-bm25-indexmsmarco-bm25
MS MARCO with hard negatives from bm25
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:
msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-bm25.SWE-bench_bm25_13K
Dataset Card for "SWE-bench_bm25_13K"
Dataset Summary
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
This dataset SWE-bench_bm25_13K includes a formatting of… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_13K.SWE-bench_bm25_40K
Dataset Card for "SWE-bench_bm25_40K"
Dataset Summary
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
This dataset SWE-bench_bm25_40K includes a formatting of… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_40K.SWE-bench_bm25_27K
Dataset Card for "SWE-bench_bm25_27K"
Dataset Summary
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
This dataset SWE-bench_bm25_27K includes a formatting of… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_bm25_27K.SWT-bench_Lite_bm25_27k_zsb
Dataset Summary
SWT-bench Lite is subset of SWT-bench, a dataset that tests systems’ ability to reproduce GitHub issues automatically. The dataset collects 276 test Issue-Pull Request pairs from 11 popular Python GitHub projects. Evaluation is performed by unit test verification using pre- and post-PR behavior of the test suite with and without the model proposed tests.
📊🏆 Leaderboard
A public leaderboard for performance on SWT-bench is hosted at swtbench.com
The… See the full description on the dataset page: https://huggingface.co/datasets/eth-sri/SWT-bench_Lite_bm25_27k_zsb.
