mteb/Core17InstructionRetrieval
Core17InstructionRetrieval An MTEB dataset Massive Text Embedding Benchmark Measuring retrieval instruction following ability on Core17 narratives for the FollowIR benchmark. Task category t2t Domains News, Written Reference https://arxiv.org/abs/2403.15246 Source datasets: jhu-clsp/core17-instructions-mteb How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/Core17InstructionRetrieval.
08.2k
1---2annotations_creators:3- derived4language:5- eng6license: mit7multilinguality: monolingual8source_datasets:9- jhu-clsp/core17-instructions-mteb10task_categories:11- text-ranking12task_ids: []13dataset_info:14- config_name: corpus15 features:16 - name: id17 dtype: string18 - name: title19 dtype: string20 - name: text21 dtype: string22 splits:23 - name: test24 num_bytes: 4484380425 num_examples: 1989926 download_size: 2783731027 dataset_size: 4484380428- config_name: instruction29 features:30 - name: query-id31 dtype: string32 - name: instruction33 dtype: string34 splits:35 - name: test36 num_bytes: 1367537 num_examples: 4038 download_size: 744339 dataset_size: 1367540- config_name: qrel_diff41 features:42 - name: query-id43 dtype: string44 - name: corpus-ids45 list: string46 splits:47 - name: qrel_diff48 num_bytes: 563249 num_examples: 2050 download_size: 564651 dataset_size: 563252- config_name: qrels53 features:54 - name: query-id55 dtype: string56 - name: corpus-id57 dtype: string58 - name: score59 dtype: int6460 splits:61 - name: test62 num_bytes: 31198063 num_examples: 948064 download_size: 5343665 dataset_size: 31198066- config_name: queries67 features:68 - name: id69 dtype: string70 - name: text71 dtype: string72 - name: instruction73 dtype: string74 splits:75 - name: test76 num_bytes: 1822577 num_examples: 4078 download_size: 1030079 dataset_size: 1822580- config_name: top_ranked81 features:82 - name: query-id83 dtype: string84 - name: corpus-ids85 list: string86 splits:87 - name: test88 num_bytes: 49850089 num_examples: 4090 download_size: 21030391 dataset_size: 49850092configs:93- config_name: corpus94 data_files:95 - split: test96 path: corpus/test-*97- config_name: instruction98 data_files:99 - split: test100 path: instruction/test-*101- config_name: qrel_diff102 data_files:103 - split: qrel_diff104 path: qrel_diff/qrel_diff-*105- config_name: qrels106 data_files:107 - split: test108 path: qrels/test-*109- config_name: queries110 data_files:111 - split: test112 path: queries/test-*113- config_name: top_ranked114 data_files:115 - split: test116 path: top_ranked/test-*117tags:118- mteb119- text120---121<!-- adapted from https://github.com/huggingface/huggingface_hub/blob/v0.30.2/src/huggingface_hub/templates/datasetcard_template.md -->122 123<div align="center" style="padding: 40px 20px; background-color: white; border-radius: 12px; box-shadow: 0 2px 10px rgba(0, 0, 0, 0.05); max-width: 600px; margin: 0 auto;">124 <h1 style="font-size: 3.5rem; color: #1a1a1a; margin: 0 0 20px 0; letter-spacing: 2px; font-weight: 700;">Core17InstructionRetrieval</h1>125 <div style="font-size: 1.5rem; color: #4a4a4a; margin-bottom: 5px; font-weight: 300;">An <a href="https://github.com/embeddings-benchmark/mteb" style="color: #2c5282; font-weight: 600; text-decoration: none;" onmouseover="this.style.textDecoration='underline'" onmouseout="this.style.textDecoration='none'">MTEB</a> dataset</div>126 <div style="font-size: 0.9rem; color: #2c5282; margin-top: 10px;">Massive Text Embedding Benchmark</div>127</div>128 129Measuring retrieval instruction following ability on Core17 narratives for the FollowIR benchmark.130 131| | |132|---------------|---------------------------------------------|133| Task category | t2t |134| Domains | News, Written |135| Reference | https://arxiv.org/abs/2403.15246 |136 137Source datasets:138- [jhu-clsp/core17-instructions-mteb](https://huggingface.co/datasets/jhu-clsp/core17-instructions-mteb)139 140 141## How to evaluate on this task142 143You can evaluate an embedding model on this dataset using the following code:144 145```python146import mteb147 148task = mteb.get_task("Core17InstructionRetrieval")149evaluator = mteb.MTEB([task])150 151model = mteb.get_model(YOUR_MODEL)152evaluator.run(model)153```154 155<!-- Datasets want link to arxiv in readme to autolink dataset with paper -->156To learn more about how to run models on `mteb` task check out the [GitHub repository](https://github.com/embeddings-benchmark/mteb).157 158## Citation159 160If you use this dataset, please cite the dataset as well as [mteb](https://github.com/embeddings-benchmark/mteb), as this dataset likely includes additional processing as a part of the [MMTEB Contribution](https://github.com/embeddings-benchmark/mteb/tree/main/docs/mmteb).161 162```bibtex163 164@misc{weller2024followir,165 archiveprefix = {arXiv},166 author = {Orion Weller and Benjamin Chang and Sean MacAvaney and Kyle Lo and Arman Cohan and Benjamin Van Durme and Dawn Lawrie and Luca Soldaini},167 eprint = {2403.15246},168 primaryclass = {cs.IR},169 title = {FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions},170 year = {2024},171}172 173 174@article{enevoldsen2025mmtebmassivemultilingualtext,175 title={MMTEB: Massive Multilingual Text Embedding Benchmark},176 author={Kenneth Enevoldsen and Isaac Chung and Imene Kerboua and Márton Kardos and Ashwin Mathur and David Stap and Jay Gala and Wissam Siblini and Dominik Krzemiński and Genta Indra Winata and Saba Sturua and Saiteja Utpala and Mathieu Ciancone and Marion Schaeffer and Gabriel Sequeira and Diganta Misra and Shreeya Dhakal and Jonathan Rystrøm and Roman Solomatin and Ömer Çağatan and Akash Kundu and Martin Bernstorff and Shitao Xiao and Akshita Sukhlecha and Bhavish Pahwa and Rafał Poświata and Kranthi Kiran GV and Shawon Ashraf and Daniel Auras and Björn Plüster and Jan Philipp Harries and Loïc Magne and Isabelle Mohr and Mariya Hendriksen and Dawei Zhu and Hippolyte Gisserot-Boukhlef and Tom Aarsen and Jan Kostkan and Konrad Wojtasik and Taemin Lee and Marek Šuppa and Crystina Zhang and Roberta Rocca and Mohammed Hamdy and Andrianos Michail and John Yang and Manuel Faysse and Aleksei Vatolin and Nandan Thakur and Manan Dey and Dipam Vasani and Pranjal Chitale and Simone Tedeschi and Nguyen Tai and Artem Snegirev and Michael Günther and Mengzhou Xia and Weijia Shi and Xing Han Lù and Jordan Clive and Gayatri Krishnakumar and Anna Maksimova and Silvan Wehrli and Maria Tikhonova and Henil Panchal and Aleksandr Abramov and Malte Ostendorff and Zheng Liu and Simon Clematide and Lester James Miranda and Alena Fenogenova and Guangyu Song and Ruqiya Bin Safi and Wen-Ding Li and Alessia Borghini and Federico Cassano and Hongjin Su and Jimmy Lin and Howard Yen and Lasse Hansen and Sara Hooker and Chenghao Xiao and Vaibhav Adlakha and Orion Weller and Siva Reddy and Niklas Muennighoff},177 publisher = {arXiv},178 journal={arXiv preprint arXiv:2502.13595},179 year={2025},180 url={https://arxiv.org/abs/2502.13595},181 doi = {10.48550/arXiv.2502.13595},182}183 184@article{muennighoff2022mteb,185 author = {Muennighoff, Niklas and Tazi, Nouamane and Magne, Loïc and Reimers, Nils},186 title = {MTEB: Massive Text Embedding Benchmark},187 publisher = {arXiv},188 journal={arXiv preprint arXiv:2210.07316},189 year = {2022}190 url = {https://arxiv.org/abs/2210.07316},191 doi = {10.48550/ARXIV.2210.07316},192}193```194 195# Dataset Statistics196<details>197 <summary> Dataset Statistics</summary>198 199The following code contains the descriptive statistics from the task. These can also be obtained using:200 201```python202import mteb203 204task = mteb.get_task("Core17InstructionRetrieval")205 206desc_stats = task.metadata.descriptive_stats207```208 209```json210{211 "test": {212 "num_samples": 19939,213 "number_of_characters": 44471883,214 "documents_text_statistics": {215 "total_text_length": 44454438,216 "min_text_length": 7,217 "average_text_length": 2234.003618272275,218 "max_text_length": 2960,219 "unique_texts": 19143220 },221 "documents_image_statistics": null,222 "queries_text_statistics": {223 "total_text_length": 17445,224 "min_text_length": 198,225 "average_text_length": 436.125,226 "max_text_length": 1000,227 "unique_texts": 40228 },229 "queries_image_statistics": null,230 "relevant_docs_statistics": {231 "num_relevant_docs": 1744,232 "min_relevant_docs_per_query": 135,233 "average_relevant_docs_per_query": 43.6,234 "max_relevant_docs_per_query": 379,235 "unique_relevant_docs": 4739236 },237 "top_ranked_statistics": {238 "num_top_ranked": 40000,239 "min_top_ranked_per_query": 1000,240 "average_top_ranked_per_query": 1000.0,241 "max_top_ranked_per_query": 1000242 }243 }244}245```246 247</details>248 249---250*This dataset card was automatically generated using [MTEB](https://github.com/embeddings-benchmark/mteb)*