datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cqadupstack-webmasters
CQADupstackWebmastersRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Benchmark Data Set for Community Question-Answering Research
Task category
t2t
Domains
Written, Web
Reference
http://nlp.cis.unimelb.edu.au/resources/cqadupstack/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackWebmastersRetrieval"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-webmasters.cqadupstack-webmasters-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackWebmasters-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-webmasters-vn.CQADupstack-Webmasters-PL
CQADupstack-Webmasters-PL
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset
Task category
t2t
Domains
Written, Web
Reference
https://huggingface.co/datasets/clarin-knext/cqadupstack-webmasters-pl
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstack-Webmasters-PL"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Webmasters-PL.beir-cqadupstack-webmasters
CQADupstackWebmastersRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackWebmastersRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-webmasters @ 160c094312a0 (the revision… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-webmasters.cqadupstack-webmasters-fa
Dataset Summary
CQADupstack-webmasters-Fa is a Persian (Farsi) dataset created for the Retrieval task, focusing on identifying duplicate or semantically similar questions within community question-answering (CQA) platforms. It is a translated version of the Webmasters StackExchange data from the English CQADupstack dataset and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark).
Language(s): Persian (Farsi)
Task(s): Retrieval (Duplicate Question Retrieval)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-webmasters-fa.beir_cqadupstack_webmasters
CQADupStack Webmasters (BEIR) — duplicate-question retrieval
Dataset description
CQADupStack is a benchmark for community question answering (cQA) built from publicly available Stack Exchange content. It was introduced by Hoogeveen, Verspoor, and Baldwin at ADCS 2015 as a resource for studying duplicate questions: threads are organized so that systems can be trained and evaluated on finding prior questions that match (or semantically duplicate) a newly asked… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/beir_cqadupstack_webmasters.cqadupstack-webmasters
Dataset Card for "cqadupstack-webmasters"
More Information needed
cqadupstack-webmasters-top-20-gen-queries
NFCorpus: 20 generated queries (BEIR Benchmark)
This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset.
DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1
id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl).
Questions generated: 20
Code used for generation: evaluate_anserini_docT5query_parallel.py
Below contains the old dataset card for the BEIR benchmark.
Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-webmasters-top-20-gen-queries.CQADupstackWebmasters-NLbeir_cqadupstack_webmasters_test
beir_cqadupstack_webmasters_test
BEIR CQADupStack/webmasters test split
Field
Value
Benchmark
beir
Sub-benchmark
cqadupstack_webmasters
Type
retrieval
Items
506
Exported from Langfuse.
CQADupstackWebmastersRetrieval-Facqadupstack-webmasters-plPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language.
Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf
Contact: konrad.wojtasik@pwr.edu.pl
cqadupstack-webmasters-pl-qrelsPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language.
Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf
Contact: konrad.wojtasik@pwr.edu.pl
cqadupstack-webmasters-qrels
Dataset Card for "cqadupstack-webmasters-qrels"
More Information needed
