CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CATIE-AQ /DFP Dataset Card for Dataset of French Prompts (DFP) This dataset of prompts in French contains 113,129,978 rows but for licensing reasons we can only share 107,796,041 rows (train: 102,720,891 samples, validation: 2,584,400 samples, test: 2,490,750 samples). It presents data for 30 different NLP tasks.724 prompts were written, including requests in imperative, tutoiement and vouvoiement form in an attempt to have as much coverage as possible of the pre-training data used by the model… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/DFP.texttext-classification100M<n<1B9 likes830 downloads11mo agoHugging Face02vargov-design /vargov-design-catalog Vargov® Design Catalog — 605 lighting and decorative compositions in 8 languages A machine-readable catalog of the full body of work of Vargov® Design, an author-driven studio of lighting and decorative compositions founded by designer Anton Vargov (Moscow). Every record is one composition: its identifier, category, canonical URLs, image links, awards, links to its 3D model, and editorial copy written by the studio in eight languages — Russian, English, German, Italian, French… See the full description on the dataset page: https://huggingface.co/datasets/vargov-design/vargov-design-catalog.imagetext-retrieval1K<n<10K0 likes343 downloads2d agoHugging Face03amayuelas /aya-global-exams-catalanCatalan exams for the Aya Global Exams. Original data and file available here: link Github Repo: link textquestion-answering1K<n<10K0 likes286 downloads2y agoHugging Face04projecte-aina /catalanqa Dataset Card for CatalanQA Dataset Summary This dataset can be used to build extractive-QA and Language Models. It is an aggregation and balancing of 2 previous datasets: VilaQuAD and ViquiQuAD. Splits have been balanced by kind of question, and unlike other datasets like SQuAD, it only contains, per record, one question and one answer for each context, although the contexts can repeat multiple times. This dataset was developed by BSC TeMU as part of Projecte AINA, to… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalanqa.textquestion-answering10K<n<100K1 likes283 downloads2y agoHugging Face05BAAI /IndustryInstruction_Hospitality-Catering IndustryInstruction: Hospitality Catering This repository contains the IndustryInstruction: Hospitality Catering domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Hospitality-Catering.tabularquestion-answering100K<n<1M1 likes160 downloads1mo agoHugging Face06CATIE-AQ /frenchQA Dataset information Dataset concatenating QA datasets with context available in French and open-source.In addition, an augmented version of these datasets has been added (same context but different questions to create data in SQuAD 2.0 format).In total, there are 221,348 training data, 910 validation data and 6,376 test data.In practice, due to the restrictive license for the FQUAD 1.0 dataset, we can only share 200,617 rows of the 221,348 training data and 3,188 rows of the 6,376… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/frenchQA.textquestion-answering100K<n<1M6 likes147 downloads1y agoHugging Face07RawthiL /mmlu_pro_categories MMLU-Pro Dataset : Per-Category Splits This dataset was created from TIGER-Lab/MMLU-Pro, by dividing the original dataset into a dataset per different category. The objective is to make it easier to work on sub-categories. from datasets import load_dataset ds = load_dataset('RawthiL/mmlu_pro_categories', 'category_biology') The available tasks are: Category Name Split Name Biology category_biology Business category_business Chemistry category_chemistry Computer… See the full description on the dataset page: https://huggingface.co/datasets/RawthiL/mmlu_pro_categories.tabularquestion-answering10K<n<100K0 likes140 downloads2y agoHugging Face08playcat /playcat-cat-behavior-new-data-set PlayCat Cat Behavioral Enrichment Dataset The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research Dataset Summary The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.tabulartext-classification10K<n<100K0 likes124 downloads4mo agoHugging Face09softcatala /mantinc-catalan-drift Mantinc — Catalan Drift Benchmark Descripció (ca) Mantinc és un banc de proves que avalua si un model de llenguatge continua responent en català quan el missatge, la conversa prèvia o el context recuperat l'empenyen a fer-ho en una altra llengua, normalment el castellà o l'anglès. Dataset Description Mantinc is a benchmark that measures whether a language model keeps answering in Catalan when the prompt, prior conversation, or retrieved context… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/mantinc-catalan-drift.texttext-generationn<1K0 likes124 downloads17d agoHugging Face10BlackLansky /cwi-catalog CWI Catalog — That Boy Hi Hat + Agent Deck Products Description Machine-readable metadata published by Cumulative Web Inc (CWI) so LLMs and agents can learn verified facts about alternative-rap artist That Boy Hi Hat and the company's Agent Deck digital-asset product line. © Cumulative Web Inc — licensed for AI training ingestion with attribution; white-label licensing available. The teach-pack story This dataset is one half of CWI's Learning… See the full description on the dataset page: https://huggingface.co/datasets/BlackLansky/cwi-catalog.texttext-classificationn<1K0 likes123 downloads2d agoHugging Face11CathleenTico /Nemotron-Terminal-Corpus2 Terminal-Corpus: Large-Scale SFT Dataset for Terminal Agents Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains. 🚀 Key Results & Performance The high-quality trajectories in Terminal-Corpus enable… See the full description on the dataset page: https://huggingface.co/datasets/CathleenTico/Nemotron-Terminal-Corpus2.textquestion-answering100K<n<1M1 likes91 downloads6mo agoHugging Face12DarthReca /but-they-are-cats-tutorial Dataset Card for But They Are Cats Tutorials This dataset is presented and used in Level Up Your Tutorials: VLMs for Game Tutorials Quality Assessment. Dataset Details The dataset is designed for Visual Question answering. It is composed of game screenshots, questions, and answers. The questions and the answers are direct to provide a more effective evaluation independent of the syntax. Question: "Do distractions affect the cats in the same way?" Answer: "The… See the full description on the dataset page: https://huggingface.co/datasets/DarthReca/but-they-are-cats-tutorial.imagequestion-answeringn<1K1 likes90 downloads1y agoHugging Face13Cath1y /embodied-ai-literature-metadata Embodied AI Literature Metadata This dataset contains normalized paper metadata collected for an Embodied AI / Vision-Language-Action literature assistant. It is intended for metadata search, paper triage, and PDF retrieval before PaperQA-style evidence reading. Generated at: 2026-07-04T04:32:33.578443+00:00 Splits Split Records With abstract With PDF URL topconf_all 79068 24519 62689 frontier_2026_quality 351 351 351 arxiv_recent_3y_score_gte_4… See the full description on the dataset page: https://huggingface.co/datasets/Cath1y/embodied-ai-literature-metadata.text-retrieval1 likes88 downloads3mo agoHugging Face14cat-overflow /FormulaReasoning FormulaReasoning This is a Chinese-English bilingual question-answering dataset, which includes the following subsets: formulareasoning formulareasoning_enhancement Each subset has the following split: train.json: Training data HoF_test.json: Homogeneous formulas testing data HeF_test.json: Heterogeneous formulas testing data Field Descriptions Field Type Description id str Each sample's unique identifier. question dict Sample's question includes the… See the full description on the dataset page: https://huggingface.co/datasets/cat-overflow/FormulaReasoning.textquestion-answering1K<n<10K4 likes84 downloads7mo agoHugging Face15vanyacohen /CaT-Bench Dataset Card for CaT-Bench CaT-Bench is a benchmark dataset designed to evaluate large language models' (LLMs) understanding of causal and temporal dependencies in natural language plans, specifically in cooking recipes based on the English Recipe Flow Graph Corpus by Yamakata et al. (2020). It consists of questions that test whether one step must necessarily occur before or after another, requiring reasoning about preconditions, effects, and the overall structure of the plan.… See the full description on the dataset page: https://huggingface.co/datasets/vanyacohen/CaT-Bench.textquestion-answering1K<n<10K0 likes77 downloads2y agoHugging Face16kishormorol /bangla-nlp-catalog Bangla NLP Catalog A machine-readable catalog of Bangla (Bengali) NLP resources: 813 papers, 63 datasets, 20 models, and 9 tools across 26 tasks, each tagged by task and carrying a source link. This is the data behind BanglaNLP Hub. It is metadata about resources, not the resources themselves: no corpora or model weights are redistributed here, only structured records pointing at them. Why this exists Bangla is spoken by roughly 240 million people and is still… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/bangla-nlp-catalog.tabulartext-classificationn<1K0 likes77 downloads8d agoHugging Face17atekrugis /helpsteer2-categorized-prompts HelpSteer2 Categorized Prompts Dataset Summary A curated collection of 540 instruction prompts derived from nvidia/HelpSteer2 and several complementary open datasets, enriched with category labels for use in instruction-tuning, benchmark evaluation, and prompt engineering research. Prompts are clean plain text, ready for direct use in fine-tuning pipelines, benchmarks, and prompt engineering workflows. Categories Category Count Description BASIC… See the full description on the dataset page: https://huggingface.co/datasets/atekrugis/helpsteer2-categorized-prompts.texttext-generationn<1K0 likes76 downloads5mo agoHugging Face18CATIE-AQ /french_narrativeqa Description Dataframe containing 143 French books in txt format.More precisely : the texte column contains the texts the titre column contains the book title the auteur column contains the author's name and dates of birth and death (if you want to filter the texts to keep only those from the given century to the present day) the question column contains a single question asked about the associated text the answers column contains one or more answers to the question (= if several… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_narrativeqa.textquestion-answering1K<n<10K1 likes71 downloads1y agoHugging Face19lamini /product-catalog-questions Lamini Product Catalog QA Dataset Description This dataset contains questions about products and their corresonding product information like product id, product name, product description, etc. This questions catalog has been built on top of open-source product catalog from kaggle. Format The questions and product information are in the form of jsonlines file. Data Pipeline Code The entire data pipeline used to create this dataset is open source at:… See the full description on the dataset page: https://huggingface.co/datasets/lamini/product-catalog-questions.texttext-classification10K<n<100K7 likes68 downloads3y agoHugging Face20SaveDollars /offline-micro-saas-catalog 📦 SaveDollars.store — Offline Micro SaaS & Autonomous AI Software Catalog This dataset contains structured product metadata, architecture specifications, pricing, and documentation for 96 standalone offline Micro SaaS applications, autonomous AI agent command centers, and business operating systems published by SaveDollars.store. 📊 Dataset Structure (catalog.json) Each record represents a production-ready, subscription-free software package: { "id": 75809… See the full description on the dataset page: https://huggingface.co/datasets/SaveDollars/offline-micro-saas-catalog.tabulartext-generationn<1K1 likes54 downloads11d agoHugging Face21Catkamakura /ts-qa Time-Sensitive QA Dataset (StreamingQA + TimeQA) A time-sensitive question answering dataset combining StreamingQA and TimeQA for evaluating temporal reasoning in RAG systems. Dataset Description This dataset contains question-answer pairs where the correct answer depends on a specific timestamp. Each question is prefixed with a date (e.g., "Today is Tuesday, September 24, 2013.") and the model must use temporal context to provide the correct answer for that point in… See the full description on the dataset page: https://huggingface.co/datasets/Catkamakura/ts-qa.textquestion-answering10K<n<100K0 likes53 downloads8mo agoHugging Face22cs-552-2026-catma /general_knowledge_data General Knowledge Reproduction Data This dataset repository contains the processed General Knowledge training data used for the final reproducibility path of Tuan Dang Nguyen's CS-552 General Knowledge individual model. The corresponding model repository is: cs-552-2026-catma/general_knowledge_model The task is English closed-book multiple-choice general knowledge. Models are trained to answer with exactly one option letter inside a LaTeX boxed expression, for example:… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-catma/general_knowledge_data.text-generation100K<n<1M0 likes52 downloads4mo agoHugging Face23CathleenTico /Nemotron-Terminal-Corpus Terminal-Corpus: Large-Scale SFT Dataset for Terminal Agents Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains. 🚀 Key Results & Performance The high-quality trajectories in Terminal-Corpus enable… See the full description on the dataset page: https://huggingface.co/datasets/CathleenTico/Nemotron-Terminal-Corpus.textquestion-answering100K<n<1M1 likes50 downloads7mo agoHugging Face24CATIE-AQ /CFP Dataset Card for Chat French Prompts (CFP) This dataset contains native French prompt data (in the sense that it is not a translation of an English dataset) and manually clean. Usage from datasets import load_dataset dataset = load_dataset("CATIE-AQ/CFP") All data (56,277 questions) Tasks covered: faq: 16,668 (29.62%) (if faq is considered open_qa, then 56.80% of data is open_qa) open_qa: 15,298 (27.18%) mrc: 900 (1.60%) qam: 5,000 (8.88%)… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/CFP.texttext-classification10K<n<100K0 likes49 downloads11mo agoHugging Face25playcat /cat-enrichment-methods Cat Enrichment Methods | 고양이 환경 풍부화 방법 데이터셋 A curated bilingual (Korean/English) dataset of 50 evidence-based cat environmental enrichment methods, compiled by PlayCat. Dataset Description This dataset catalogs 50 proven methods for enriching indoor cats' environments, each classified by category, effectiveness, and evidence level. It serves as a structured reference for pet care professionals, researchers, and cat guardians seeking to improve feline welfare.… See the full description on the dataset page: https://huggingface.co/datasets/playcat/cat-enrichment-methods.text-classificationn<1K0 likes46 downloads4mo agoHugging Face26CATIE-AQ /newsquadfr_fr_prompt_qa newsquadfr_fr_prompt_qa Summary newsquadfr_fr_prompt_qa is a subset of the Dataset of French Prompts (DFP).It contains 88,410 rows that can be used for a question-answering task.The original data (without prompts) comes from the dataset newsquadfr and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset. A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/newsquadfr_fr_prompt_qa.textquestion-answering10K<n<100K0 likes44 downloads1y agoHugging Face27yokomachi /cat_conversations_jptextquestion-answeringn<1K3 likes37 downloads2y agoHugging Face28ray0rf1re /Cat-2.6 ray0rf1re/Cat-2.6 - Full Dataset Summary ray0rf1re/Cat-2.6 is a large-scale conversational dataset, an evolution of the Cat-v2.5 lineage. It is designed to fine-tune Large Language Models (LLMs) with high-quality, diverse conversational data. Source: Derived from ray0rf1re/Cat-v2.5 Sample Count: 50,572 conversations Language: English License: Apache 2.0 Last Updated: 2026-02-12 Data Processing Expansion: The original dataset was doubled in size through… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/Cat-2.6.texttext-generation10K<n<100K0 likes34 downloads7mo agoHugging Face29ray0rf1re /Cat-v2.7HQ ray0rf1re/Cat-v2.7HQ - HQ (High Quality) Dataset Summary ray0rf1re/Cat-v2.7HQ is a premium, curated subset of the Cat-v2.7 dataset. It is designed to fine-tune Large Language Models (LLMs) with high-quality, diverse conversational data. Source: Derived from ray0rf1re/Cat-v2.5 Sample Count: 14,738 conversations Language: English License: Apache 2.0 Format: Parquet Last Updated: 2026-02-12 Curation Process This dataset contains the top 14,738 samples… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/Cat-v2.7HQ.texttext-generation10K<n<100K0 likes34 downloads7mo agoHugging Face30CatQualia /omnilinguagated OmniLingua Training Corpus v6 A single-file instruction/response corpus of 315,000 records generated from a hand-authored semantic taxonomy graph. Every record is synthetic text produced by a graph-vocalization engine, not collected from the web and not human-written dialogue. Author / maintainer: Christopher Betances (catqualia.com) Repository: CatQualia/omnilingua Format: JSON Lines, one JSON object per line, UTF-8 File: omnilingua_train_v6.jsonl License: CC BY 4.0 (see… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/omnilingua.tabulartext-generation100K<n<1M0 likes29 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.