CoolFace
Datasetpublic

mteb/cqadupstack-android

CQADupstackAndroidRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Programming, Web, Written, Non-fiction Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-android.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes430downloads
README.md197 linesDownload Raw Back to root
1---2annotations_creators:3- derived4language:5- eng6license: apache-2.07multilinguality: monolingual8task_categories:9- text-retrieval10task_ids:11- multiple-choice-qa12config_names:13- corpus14tags:15- mteb16- text17dataset_info:18- config_name: default19  features:20  - name: query-id21    dtype: string22  - name: corpus-id23    dtype: string24  - name: score25    dtype: float6426  splits:27  - name: test28    num_bytes: 4341129    num_examples: 169630- config_name: corpus31  features:32  - name: _id33    dtype: string34  - name: title35    dtype: string36  - name: text37    dtype: string38  splits:39  - name: corpus40    num_bytes: 1404446941    num_examples: 2299842- config_name: queries43  features:44  - name: _id45    dtype: string46  - name: text47    dtype: string48  splits:49  - name: queries50    num_bytes: 4515751    num_examples: 69952configs:53- config_name: default54  data_files:55  - split: test56    path: qrels/test.jsonl57- config_name: corpus58  data_files:59  - split: corpus60    path: corpus.jsonl61- config_name: queries62  data_files:63  - split: queries64    path: queries.jsonl65---66<!-- adapted from https://github.com/huggingface/huggingface_hub/blob/v0.30.2/src/huggingface_hub/templates/datasetcard_template.md -->67 68<div align="center" style="padding: 40px 20px; background-color: white; border-radius: 12px; box-shadow: 0 2px 10px rgba(0, 0, 0, 0.05); max-width: 600px; margin: 0 auto;">69  <h1 style="font-size: 3.5rem; color: #1a1a1a; margin: 0 0 20px 0; letter-spacing: 2px; font-weight: 700;">CQADupstackAndroidRetrieval</h1>70  <div style="font-size: 1.5rem; color: #4a4a4a; margin-bottom: 5px; font-weight: 300;">An <a href="https://github.com/embeddings-benchmark/mteb" style="color: #2c5282; font-weight: 600; text-decoration: none;" onmouseover="this.style.textDecoration='underline'" onmouseout="this.style.textDecoration='none'">MTEB</a> dataset</div>71  <div style="font-size: 0.9rem; color: #2c5282; margin-top: 10px;">Massive Text Embedding Benchmark</div>72</div>73 74CQADupStack: A Benchmark Data Set for Community Question-Answering Research75 76|               |                                             |77|---------------|---------------------------------------------|78| Task category | t2t                              |79| Domains       | Programming, Web, Written, Non-fiction                               |80| Reference     | http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ |81 82 83## How to evaluate on this task84 85You can evaluate an embedding model on this dataset using the following code:86 87```python88import mteb89 90task = mteb.get_tasks(["CQADupstackAndroidRetrieval"])91evaluator = mteb.MTEB(task)92 93model = mteb.get_model(YOUR_MODEL)94evaluator.run(model)95```96 97<!-- Datasets want link to arxiv in readme to autolink dataset with paper -->98To learn more about how to run models on `mteb` task check out the [GitHub repitory](https://github.com/embeddings-benchmark/mteb). 99 100## Citation101 102If you use this dataset, please cite the dataset as well as [mteb](https://github.com/embeddings-benchmark/mteb), as this dataset likely includes additional processing as a part of the [MMTEB Contribution](https://github.com/embeddings-benchmark/mteb/tree/main/docs/mmteb).103 104```bibtex105 106@inproceedings{hoogeveen2015,107  acmid = {2838934},108  address = {New York, NY, USA},109  articleno = {3},110  author = {Hoogeveen, Doris and Verspoor, Karin M. and Baldwin, Timothy},111  booktitle = {Proceedings of the 20th Australasian Document Computing Symposium (ADCS)},112  doi = {10.1145/2838931.2838934},113  isbn = {978-1-4503-4040-3},114  location = {Parramatta, NSW, Australia},115  numpages = {8},116  pages = {3:1--3:8},117  publisher = {ACM},118  series = {ADCS '15},119  title = {CQADupStack: A Benchmark Data Set for Community Question-Answering Research},120  url = {http://doi.acm.org/10.1145/2838931.2838934},121  year = {2015},122}123 124 125@article{enevoldsen2025mmtebmassivemultilingualtext,126  title={MMTEB: Massive Multilingual Text Embedding Benchmark},127  author={Kenneth Enevoldsen and Isaac Chung and Imene Kerboua and Márton Kardos and Ashwin Mathur and David Stap and Jay Gala and Wissam Siblini and Dominik Krzemiński and Genta Indra Winata and Saba Sturua and Saiteja Utpala and Mathieu Ciancone and Marion Schaeffer and Gabriel Sequeira and Diganta Misra and Shreeya Dhakal and Jonathan Rystrøm and Roman Solomatin and Ömer Çağatan and Akash Kundu and Martin Bernstorff and Shitao Xiao and Akshita Sukhlecha and Bhavish Pahwa and Rafał Poświata and Kranthi Kiran GV and Shawon Ashraf and Daniel Auras and Björn Plüster and Jan Philipp Harries and Loïc Magne and Isabelle Mohr and Mariya Hendriksen and Dawei Zhu and Hippolyte Gisserot-Boukhlef and Tom Aarsen and Jan Kostkan and Konrad Wojtasik and Taemin Lee and Marek Šuppa and Crystina Zhang and Roberta Rocca and Mohammed Hamdy and Andrianos Michail and John Yang and Manuel Faysse and Aleksei Vatolin and Nandan Thakur and Manan Dey and Dipam Vasani and Pranjal Chitale and Simone Tedeschi and Nguyen Tai and Artem Snegirev and Michael Günther and Mengzhou Xia and Weijia Shi and Xing Han Lù and Jordan Clive and Gayatri Krishnakumar and Anna Maksimova and Silvan Wehrli and Maria Tikhonova and Henil Panchal and Aleksandr Abramov and Malte Ostendorff and Zheng Liu and Simon Clematide and Lester James Miranda and Alena Fenogenova and Guangyu Song and Ruqiya Bin Safi and Wen-Ding Li and Alessia Borghini and Federico Cassano and Hongjin Su and Jimmy Lin and Howard Yen and Lasse Hansen and Sara Hooker and Chenghao Xiao and Vaibhav Adlakha and Orion Weller and Siva Reddy and Niklas Muennighoff},128  publisher = {arXiv},129  journal={arXiv preprint arXiv:2502.13595},130  year={2025},131  url={https://arxiv.org/abs/2502.13595},132  doi = {10.48550/arXiv.2502.13595},133}134 135@article{muennighoff2022mteb,136  author = {Muennighoff, Niklas and Tazi, Nouamane and Magne, Lo{\"\i}c and Reimers, Nils},137  title = {MTEB: Massive Text Embedding Benchmark},138  publisher = {arXiv},139  journal={arXiv preprint arXiv:2210.07316},140  year = {2022}141  url = {https://arxiv.org/abs/2210.07316},142  doi = {10.48550/ARXIV.2210.07316},143}144```145 146# Dataset Statistics147<details>148  <summary> Dataset Statistics</summary>149 150The following code contains the descriptive statistics from the task. These can also be obtained using:151 152```python153import mteb154 155task = mteb.get_task("CQADupstackAndroidRetrieval")156 157desc_stats = task.metadata.descriptive_stats158```159 160```json161{162    "test": {163        "num_samples": 23697,164        "number_of_characters": 13713141,165        "num_documents": 22998,166        "min_document_length": 57,167        "average_document_length": 594.701974084703,168        "max_document_length": 27831,169        "unique_documents": 22998,170        "num_queries": 699,171        "min_query_length": 16,172        "average_query_length": 51.76680972818312,173        "max_query_length": 127,174        "unique_queries": 699,175        "none_queries": 0,176        "num_relevant_docs": 1696,177        "min_relevant_docs_per_query": 1,178        "average_relevant_docs_per_query": 2.4263233190271816,179        "max_relevant_docs_per_query": 262,180        "unique_relevant_docs": 1696,181        "num_instructions": null,182        "min_instruction_length": null,183        "average_instruction_length": null,184        "max_instruction_length": null,185        "unique_instructions": null,186        "num_top_ranked": null,187        "min_top_ranked_per_query": null,188        "average_top_ranked_per_query": null,189        "max_top_ranked_per_query": null190    }191}192```193 194</details>195 196---197*This dataset card was automatically generated using [MTEB](https://github.com/embeddings-benchmark/mteb)*