CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /cqadupstack-gaming CQADupstackGamingRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Web, Written Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackGamingRetrieval"]) evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-gaming.texttext-retrieval10K<n<100K0 likes6.4k downloads1y agoHugging Face02mteb /cqadupstack-unix CQADupstackUnixRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Web, Programming Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackUnixRetrieval"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-unix.texttext-retrieval10K<n<100K0 likes6.3k downloads1y agoHugging Face03clips /beir-nl-cqadupstack Dataset Card for BEIR-NL Benchmark Dataset Summary BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB). BEIR-NL contains the following tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-cqadupstack.texttext-retrieval100K<n<1M0 likes5.6k downloads2y agoHugging Face04mteb /cqadupstack-physics CQADupstackPhysicsRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Academic, Non-fiction Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackPhysicsRetrieval"]) evaluator… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-physics.texttext-retrieval10K<n<100K2 likes1.7k downloads1y agoHugging Face05mteb /cqadupstack-english CQADupstackEnglishRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackEnglishRetrieval"]) evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-english.texttext-retrieval10K<n<100K1 likes1.6k downloads1y agoHugging Face06mteb /cqadupstack-programmers CQADupstackProgrammersRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Programming, Written, Non-fiction Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-programmers.texttext-retrieval10K<n<100K0 likes1.2k downloads1y agoHugging Face07mteb /cqadupstack-wordpress CQADupstackWordpressRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Web, Programming Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackWordpressRetrieval"]) evaluator… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-wordpress.texttext-retrieval10K<n<100K2 likes1.1k downloads1y agoHugging Face08mteb /cqadupstack-stats CQADupstackStatsRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Academic, Non-fiction Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackStatsRetrieval"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-stats.texttext-retrieval10K<n<100K0 likes1.1k downloads1y agoHugging Face09mteb /cqadupstack-gis CQADupstackGisRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Non-fiction Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackGisRetrieval"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-gis.texttext-retrieval10K<n<100K3 likes1.1k downloads1y agoHugging Face10mteb /cqadupstack-mathematica CQADupstackMathematicaRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Academic, Non-fiction Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackMathematicaRetrieval"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-mathematica.texttext-retrieval10K<n<100K1 likes1k downloads1y agoHugging Face11mteb /cqadupstack-webmasters CQADupstackWebmastersRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Web Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackWebmastersRetrieval"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-webmasters.texttext-retrieval10K<n<100K0 likes1k downloads1y agoHugging Face12mteb /cqadupstack-tex CQADupstackTexRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Non-fiction Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackTexRetrieval"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-tex.texttext-retrieval10K<n<100K1 likes987 downloads1y agoHugging Face13mteb /cqadupstack-android CQADupstackAndroidRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Programming, Web, Written, Non-fiction Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackAndroidRetrieval"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-android.texttext-retrieval10K<n<100K0 likes417 downloads1y agoHugging Face14BeIR /cqadupstack-generated-queries Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/cqadupstack-generated-queries.texttext-retrieval1M<n<10M1 likes60 downloads4y agoHugging Face15LLukas22 /cqadupstack Dataset Card for "cqadupstack" Dataset Summary This is a preprocessed version of cqadupstack, to make it easily consumable via huggingface. The original dataset can be found here. CQADupStack is a benchmark dataset for community question-answering (cQA) research. It contains threads from twelve StackExchange1 subforums, annotated with duplicate question information and comes with pre-defined training, development, and test splits, both for retrieval and classification… See the full description on the dataset page: https://huggingface.co/datasets/LLukas22/cqadupstack.textsentence-similarity100K<n<1M0 likes51 downloads3y agoHugging Face16MCINext /cqadupstack-physics-fa Dataset Summary CQADupstack-physics-Fa is a Persian (Farsi) dataset developed for the Retrieval task, with a focus on duplicate question detection in community question-answering (CQA) platforms. This dataset is a translated version of the "Physics" StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source: Translated from… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-physics-fa.text10K<n<100K0 likes47 downloads1y agoHugging Face17MCINext /cqadupstack-gaming-fa Dataset Summary CQADupstack-gaming-Fa is a Persian (Farsi) dataset developed for the Retrieval task, with a focus on duplicate question detection in community question-answering (CQA) forums. This dataset is a translated version of the "gaming" (Arqade) StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source: Translated from… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-gaming-fa.text10K<n<100K0 likes38 downloads1y agoHugging Face18income /cqadupstack-programmers-top-20-gen-queries NFCorpus: 20 generated queries (BEIR Benchmark) This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset. DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1 id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl). Questions generated: 20 Code used for generation: evaluate_anserini_docT5query_parallel.py Below contains the old dataset card for the BEIR benchmark. Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-programmers-top-20-gen-queries.texttext-retrieval10K<n<100K0 likes36 downloads4y agoHugging Face19MCINext /cqadupstack-gis-fa Dataset Summary CQADupstack-gis-Fa is a Persian (Farsi) dataset developed for the Retrieval task, with a focus on duplicate question detection in community question-answering (CQA) platforms. This dataset is a translated version of the "GIS" (Geographic Information Systems) StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-gis-fa.text10K<n<100K0 likes33 downloads1y agoHugging Face20MCINext /cqadupstack-wordpress-fa Dataset Summary CQADupstack-wordpress-Fa is a Persian (Farsi) dataset created for the Retrieval task, focused on identifying duplicate or semantically equivalent questions in the domain of WordPress development. It is a translated version of the WordPress Development StackExchange data from the English CQADupstack dataset and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-wordpress-fa.text10K<n<100K0 likes33 downloads1y agoHugging Face21MCINext /cqadupstack-mathematica-fa Dataset Summary CQADupstack-mathematica-Fa is a Persian (Farsi) dataset created for the Retrieval task, focused on duplicate question detection in community question-answering (CQA) forums. It is a translated version of the "mathematica" (Mathematica Stack Exchange) subforum from the English CQADupstack dataset and is part of the FaMTEB benchmark within the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source: Translated from… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-mathematica-fa.text10K<n<100K0 likes31 downloads1y agoHugging Face22income /cqadupstack-physics-top-20-gen-queries NFCorpus: 20 generated queries (BEIR Benchmark) This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset. DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1 id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl). Questions generated: 20 Code used for generation: evaluate_anserini_docT5query_parallel.py Below contains the old dataset card for the BEIR benchmark. Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-physics-top-20-gen-queries.texttext-retrieval10K<n<100K2 likes29 downloads4y agoHugging Face23MCINext /cqadupstack-android-fa Dataset Summary CQADupstack-android-Fa is a Persian (Farsi) dataset developed for the Retrieval task, with a focus on duplicate question detection in community question-answering (CQA) forums. This dataset is a translated version of the "android" (Android Enthusiasts) StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-android-fa.text10K<n<100K0 likes27 downloads1y agoHugging Face24MCINext /cqadupstack-webmasters-fa Dataset Summary CQADupstack-webmasters-Fa is a Persian (Farsi) dataset created for the Retrieval task, focusing on identifying duplicate or semantically similar questions within community question-answering (CQA) platforms. It is a translated version of the Webmasters StackExchange data from the English CQADupstack dataset and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-webmasters-fa.text10K<n<100K0 likes25 downloads1y agoHugging Face25income /cqadupstack-wordpress-top-20-gen-queries NFCorpus: 20 generated queries (BEIR Benchmark) This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset. DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1 id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl). Questions generated: 20 Code used for generation: evaluate_anserini_docT5query_parallel.py Below contains the old dataset card for the BEIR benchmark. Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-wordpress-top-20-gen-queries.texttext-retrieval10K<n<100K3 likes23 downloads4y agoHugging Face26MCINext /cqadupstack-english-fa Dataset Summary CQADupstack-english-Fa is a Persian (Farsi) dataset developed for the Retrieval task, focused on duplicate question detection in community question-answering (CQA) forums. This dataset is a translated version of the "english" (English Language & Usage) StackExchange subforum from the original English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-english-fa.text10K<n<100K0 likes23 downloads1y agoHugging Face27MCINext /cqadupstack-programmers-fa Dataset Summary CQADupstack-programmers-Fa is a Persian (Farsi) dataset developed for the Retrieval task, with a focus on duplicate question detection in community question-answering (CQA) platforms. This dataset is a translated version of the "Programmers" (Software Engineering) StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-programmers-fa.text10K<n<100K0 likes22 downloads1y agoHugging Face28MCINext /cqadupstack-tex-fa Dataset Summary CQADupstack-tex-Fa is a Persian (Farsi) dataset curated for the Retrieval task, specifically targeting duplicate question detection in community question-answering (CQA) forums. This dataset is a translation of the "TeX - LaTeX" StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source: Translated from English… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-tex-fa.text10K<n<100K0 likes21 downloads1y agoHugging Face29kaengreg /rus-cqadupstacktext100K<n<1M0 likes21 downloads2y agoHugging Face30MCINext /cqadupstack-stats-fa Dataset Summary CQADupstack-stats-Fa is a Persian (Farsi) dataset developed for the Retrieval task, focusing on duplicate question detection in community question-answering (CQA) platforms. This dataset is a translated version of the "Cross Validated" (Stats) subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source: Translated from English… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-stats-fa.text10K<n<100K0 likes20 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.