CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /cqadupstack-wordpress CQADupstackWordpressRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Web, Programming Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackWordpressRetrieval"]) evaluator… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-wordpress.texttext-retrieval10K<n<100K2 likes1.2k downloads1y agoHugging Face02mteb /CQADupstack-Wordpress-PL CQADupstack-Wordpress-PL An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset Task category t2t Domains Written, Web, Programming Reference https://huggingface.co/datasets/clarin-knext/cqadupstack-wordpress-pl How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstack-Wordpress-PL"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Wordpress-PL.texttext-retrieval10K<n<100K0 likes83 downloads1y agoHugging Face03Hyukkyu /beir-cqadupstack-wordpress CQADupstackWordpressRetrieval — BEIR, unified schema A normalised copy of the dataset behind the mteb task CQADupstackWordpressRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset in this collection. Source mteb/cqadupstack-wordpress @ 4ffe81d471b1 (the revision… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-wordpress.texttext-retrieval10K<n<100K0 likes47 downloads14d agoHugging Face04reapxdev /wordpress-plugins-scraper WordPress Plugins Scraper · Plugins, Installs & Ratings Scrape WordPress plugins directory by tags, search queries, author accounts, install bands, and rating filters. Extract ratings, active installs, tags, release details, and author links. Rows in this dataset 1,660 Fields 23 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/wordpress-plugins-scraper.tabular1K<n<10K0 likes40 downloads2mo agoHugging Face05GreenNode /cqadupstack-wordpress-vn How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackWordpress-VN"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL) evaluator.run(model) To learn more about how to run models on mteb task check out the GitHub repitory. Citation If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-wordpress-vn.texttext-retrieval10K<n<100K0 likes38 downloads1y agoHugging Face06alexfrancow /wordpressBytes of javascript and css files in wordpress applications for multiple classification and identification wordpress versions. tabular10K<n<100K0 likes34 downloads4y agoHugging Face07MCINext /cqadupstack-wordpress-fa Dataset Summary CQADupstack-wordpress-Fa is a Persian (Farsi) dataset created for the Retrieval task, focused on identifying duplicate or semantically equivalent questions in the domain of WordPress development. It is a translated version of the WordPress Development StackExchange data from the English CQADupstack dataset and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-wordpress-fa.text10K<n<100K0 likes34 downloads1y agoHugging Face08haohaa /shan-wordpress Language Shan - shn textn<1K0 likes29 downloads2y agoHugging Face09income /cqadupstack-wordpress-top-20-gen-queries NFCorpus: 20 generated queries (BEIR Benchmark) This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset. DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1 id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl). Questions generated: 20 Code used for generation: evaluate_anserini_docT5query_parallel.py Below contains the old dataset card for the BEIR benchmark. Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-wordpress-top-20-gen-queries.texttext-retrieval10K<n<100K3 likes25 downloads4y agoHugging Face10dmrau /cqadupstack-wordpress Dataset Card for "cqadupstack-wordpress" More Information needed text10K<n<100K1 likes20 downloads3y agoHugging Face11Karmane /wordpress-booking-appointment-plugins-market-intelligence-sample WordPress Booking & Appointment Plugins Market Intelligence Dataset -- Free Evaluation Sample This dataset packages public WordPress.org plugin-directory records for booking, appointment, and reservation plugins into one analysis-ready market-intelligence table. Each row represents one WordPress plugin enriched with install and rating metrics, update recency, support-resolution signals, commercial-language flags, booking-workflow feature flags, buyer-segment heuristics, and… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/wordpress-booking-appointment-plugins-market-intelligence-sample.imagetabular-classificationn<1K0 likes20 downloads4mo agoHugging Face12Superdav42 /wordpress-translations-nl WordPress Translation Dataset (English → NL) Translation pairs extracted from WordPress plugins, themes, and core translations. Designed for fine-tuning LLMs on WordPress-specific translation tasks. Dataset Description This dataset contains English to NL translation pairs extracted from the official WordPress translation project (translate.wordpress.org). Dataset Statistics Metric Value Training examples 222,216 Test examples 55,554 Avg source… See the full description on the dataset page: https://huggingface.co/datasets/Superdav42/wordpress-translations-nl.texttranslation100K<n<1M0 likes19 downloads9mo agoHugging Face13dmrau /cqudubstack-wordpress Dataset Card for "cqudubstack-wordpress" More Information needed text10K<n<100K1 likes16 downloads3y agoHugging Face14NODARISHUB /mx-wordpress-seo-health-benchmark WordPress/PHP Technical SEO Health Benchmark — 4 Mexican SME Sites A comparable technical-SEO health benchmark across 4 real, live production websites in Mexico, spanning different stacks: WordPress (hand-coded theme), WordPress (Astra + Elementor), WordPress (WooCommerce/Elementor), and a CMS-free vanilla PHP + MySQL site. All four are scored with the same deterministic rubric (Technical, On-Page, Speed, Headers → 0-100), so scores are directly comparable across sites and… See the full description on the dataset page: https://huggingface.co/datasets/NODARISHUB/mx-wordpress-seo-health-benchmark.tabularn<1K0 likes16 downloads2mo agoHugging Face15orgrctera /beir_cqadupstack_wordpress_test beir_cqadupstack_wordpress_test BEIR CQADupStack/wordpress test split Field Value Benchmark beir Sub-benchmark cqadupstack_wordpress Type retrieval Items 541 Exported from Langfuse. textquestion-answeringn<1K0 likes12 downloads7mo agoHugging Face16orgrctera /beir_cqadupstack_wordpress CQADupStack WordPress (BEIR) — duplicate-question retrieval Dataset description CQADupStack is a benchmark for community question answering (cQA) built from publicly available Stack Exchange content. It was introduced by Hoogeveen, Verspoor, and Baldwin at ADCS 2015 as a resource for studying duplicate questions: threads and posts are organized so that systems can be trained and evaluated on finding prior questions that match (or semantically duplicate) a newly asked… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/beir_cqadupstack_wordpress.texttext-retrievaln<1K0 likes12 downloads6mo agoHugging Face17dmrau /cqadupstack-wordpress-qrels Dataset Card for "cqadupstack-wordpress-qrels" More Information needed textn<1K0 likes9 downloads3y agoHugging Face18mattPearce /wordpress-blocks-sfttextn<1K0 likes9 downloads10mo agoHugging Face19dmrau /cqadubstack-wordpress-qrels Dataset Card for "cqadubstack-wordpress-qrels" More Information needed textn<1K0 likes8 downloads3y agoHugging Face20Karmane /wordpress-ai-mcp-plugins-enrichedgated WordPress AI and MCP Plugins Enriched Dataset This dataset packages public WordPress.org plugin-directory records for AI, LLM, chatbot, and MCP-related plugins into a single analysis-ready dataframe. Each row represents one WordPress plugin enriched with normalized descriptions, query-term provenance, install and rating metrics, provider keyword coverage, commercialization signals, MCP positioning, and category heuristics. The result is easier to use for plugin market research… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/wordpress-ai-mcp-plugins-enriched.imagetabular-classification1K<n<10K1 likes8 downloads4mo agoHugging Face21tinysec /wordpresstext1K<n<10K0 likes7 downloads1y agoHugging Face22Karmane /wordpress-booking-appointment-plugins-market-intelligencegated WordPress Booking & Appointment Plugins Market Intelligence Dataset This dataset packages public WordPress.org plugin-directory records for booking, appointment, and reservation plugins into one analysis-ready market-intelligence table. Each row represents one WordPress plugin enriched with install and rating metrics, update recency, support-resolution signals, commercial-language flags, booking-workflow feature flags, buyer-segment heuristics, and normalized source metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/wordpress-booking-appointment-plugins-market-intelligence.imagetabular-classificationn<1K0 likes4 downloads4mo agoHugging Face23prappo /wordpress-gutenberg-block-patterns WordPress Gutenberg Block Patterns texttext-generationn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.