datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cqadupstack-tex
CQADupstackTexRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Benchmark Data Set for Community Question-Answering Research
Task category
t2t
Domains
Written, Non-fiction
Reference
http://nlp.cis.unimelb.edu.au/resources/cqadupstack/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackTexRetrieval"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-tex.cqadupstack-tex-fa
Dataset Summary
CQADupstack-tex-Fa is a Persian (Farsi) dataset curated for the Retrieval task, specifically targeting duplicate question detection in community question-answering (CQA) forums. This dataset is a translation of the "TeX - LaTeX" StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite.
Language(s): Persian (Farsi)
Task(s): Retrieval (Duplicate Question Retrieval)
Source: Translated from English… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-tex-fa.cqadupstack-tex-top-20-gen-queries
NFCorpus: 20 generated queries (BEIR Benchmark)
This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset.
DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1
id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl).
Questions generated: 20
Code used for generation: evaluate_anserini_docT5query_parallel.py
Below contains the old dataset card for the BEIR benchmark.
Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-tex-top-20-gen-queries.
