datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cqadupstack-wordpress
CQADupstackWordpressRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Benchmark Data Set for Community Question-Answering Research
Task category
t2t
Domains
Written, Web, Programming
Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackWordpressRetrieval"])
evaluator… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-wordpress.wordpress-plugins-scraper
WordPress Plugins Scraper · Plugins, Installs & Ratings
Scrape WordPress plugins directory by tags, search queries, author accounts, install bands, and rating filters. Extract ratings, active installs, tags, release details, and author links.
Rows in this dataset
1,660
Fields
23
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/wordpress-plugins-scraper.cqadupstack-wordpress-fa
Dataset Summary
CQADupstack-wordpress-Fa is a Persian (Farsi) dataset created for the Retrieval task, focused on identifying duplicate or semantically equivalent questions in the domain of WordPress development. It is a translated version of the WordPress Development StackExchange data from the English CQADupstack dataset and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark).
Language(s): Persian (Farsi)
Task(s): Retrieval (Duplicate Question Retrieval)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-wordpress-fa.shan-wordpress
Language
Shan - shn
cqadupstack-wordpress-top-20-gen-queries
NFCorpus: 20 generated queries (BEIR Benchmark)
This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset.
DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1
id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl).
Questions generated: 20
Code used for generation: evaluate_anserini_docT5query_parallel.py
Below contains the old dataset card for the BEIR benchmark.
Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-wordpress-top-20-gen-queries.wordpress-blocks-sft
