datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DFP
Dataset Card for Dataset of French Prompts (DFP)
This dataset of prompts in French contains 113,129,978 rows but for licensing reasons we can only share 107,796,041 rows (train: 102,720,891 samples, validation: 2,584,400 samples, test: 2,490,750 samples). It presents data for 30 different NLP tasks.724 prompts were written, including requests in imperative, tutoiement and vouvoiement form in an attempt to have as much coverage as possible of the pre-training data used by the model… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/DFP.vargov-design-catalog
Vargov® Design Catalog — 605 lighting and decorative compositions in 8 languages
A machine-readable catalog of the full body of work of Vargov® Design, an author-driven
studio of lighting and decorative compositions founded by designer Anton Vargov (Moscow).
Every record is one composition: its identifier, category, canonical URLs, image links,
awards, links to its 3D model, and editorial copy written by the studio in eight
languages — Russian, English, German, Italian, French… See the full description on the dataset page: https://huggingface.co/datasets/vargov-design/vargov-design-catalog.aya-global-exams-catalanCatalan exams for the Aya Global Exams.
Original data and file available here: link
Github Repo: link
catalanqa
Dataset Card for CatalanQA
Dataset Summary
This dataset can be used to build extractive-QA and Language Models. It is an aggregation and balancing of 2 previous datasets: VilaQuAD and ViquiQuAD.
Splits have been balanced by kind of question, and unlike other datasets like SQuAD, it only contains, per record, one question and one answer for each context, although the contexts can repeat multiple times.
This dataset was developed by BSC TeMU as part of Projecte AINA, to… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalanqa.IndustryInstruction_Hospitality-Catering
IndustryInstruction: Hospitality Catering
This repository contains the IndustryInstruction: Hospitality Catering domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Hospitality-Catering.frenchQA
Dataset information
Dataset concatenating QA datasets with context available in French and open-source.In addition, an augmented version of these datasets has been added (same context but different questions to create data in SQuAD 2.0 format).In total, there are 221,348 training data, 910 validation data and 6,376 test data.In practice, due to the restrictive license for the FQUAD 1.0 dataset, we can only share 200,617 rows of the 221,348 training data and 3,188 rows of the 6,376… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/frenchQA.mmlu_pro_categories
MMLU-Pro Dataset : Per-Category Splits
This dataset was created from TIGER-Lab/MMLU-Pro, by dividing the original dataset into a dataset per different category. The objective is to make it easier to work on sub-categories.
from datasets import load_dataset
ds = load_dataset('RawthiL/mmlu_pro_categories', 'category_biology')
The available tasks are:
Category Name
Split Name
Biology
category_biology
Business
category_business
Chemistry
category_chemistry
Computer… See the full description on the dataset page: https://huggingface.co/datasets/RawthiL/mmlu_pro_categories.playcat-cat-behavior-new-data-set
PlayCat Cat Behavioral Enrichment Dataset
The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research
Dataset Summary
The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.mantinc-catalan-drift
Mantinc — Catalan Drift Benchmark
Descripció (ca)
Mantinc és un banc de proves que avalua si un model de llenguatge continua
responent en català quan el missatge, la conversa prèvia o el context recuperat
l'empenyen a fer-ho en una altra llengua, normalment el castellà o l'anglès.
Dataset Description
Mantinc is a benchmark that measures whether a language model keeps answering
in Catalan when the prompt, prior conversation, or retrieved context… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/mantinc-catalan-drift.cwi-catalog
CWI Catalog — That Boy Hi Hat + Agent Deck Products
Description
Machine-readable metadata published by Cumulative Web Inc (CWI) so LLMs and agents can learn verified facts about alternative-rap artist That Boy Hi Hat and the company's Agent Deck digital-asset product line.
© Cumulative Web Inc — licensed for AI training ingestion with attribution; white-label licensing available.
The teach-pack story
This dataset is one half of CWI's Learning… See the full description on the dataset page: https://huggingface.co/datasets/BlackLansky/cwi-catalog.Nemotron-Terminal-Corpus2
Terminal-Corpus: Large-Scale SFT Dataset for Terminal Agents
Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains.
🚀 Key Results & Performance
The high-quality trajectories in Terminal-Corpus enable… See the full description on the dataset page: https://huggingface.co/datasets/CathleenTico/Nemotron-Terminal-Corpus2.but-they-are-cats-tutorial
Dataset Card for But They Are Cats Tutorials
This dataset is presented and used in Level Up Your Tutorials: VLMs for Game Tutorials Quality Assessment.
Dataset Details
The dataset is designed for Visual Question answering. It is composed of game screenshots, questions, and answers.
The questions and the answers are direct to provide a more effective evaluation independent of the syntax.
Question: "Do distractions affect the cats in the same way?" Answer: "The… See the full description on the dataset page: https://huggingface.co/datasets/DarthReca/but-they-are-cats-tutorial.embodied-ai-literature-metadata
Embodied AI Literature Metadata
This dataset contains normalized paper metadata collected for an Embodied AI /
Vision-Language-Action literature assistant. It is intended for metadata search,
paper triage, and PDF retrieval before PaperQA-style evidence reading.
Generated at: 2026-07-04T04:32:33.578443+00:00
Splits
Split
Records
With abstract
With PDF URL
topconf_all
79068
24519
62689
frontier_2026_quality
351
351
351
arxiv_recent_3y_score_gte_4… See the full description on the dataset page: https://huggingface.co/datasets/Cath1y/embodied-ai-literature-metadata.FormulaReasoning
FormulaReasoning
This is a Chinese-English bilingual question-answering dataset, which includes the following subsets:
formulareasoning
formulareasoning_enhancement
Each subset has the following split:
train.json: Training data
HoF_test.json: Homogeneous formulas testing data
HeF_test.json: Heterogeneous formulas testing data
Field Descriptions
Field
Type
Description
id
str
Each sample's unique identifier.
question
dict
Sample's question includes the… See the full description on the dataset page: https://huggingface.co/datasets/cat-overflow/FormulaReasoning.CaT-Bench
Dataset Card for CaT-Bench
CaT-Bench is a benchmark dataset designed to evaluate large language models' (LLMs) understanding of causal and temporal dependencies in natural language plans, specifically in cooking recipes based on the English Recipe Flow Graph Corpus by Yamakata et al. (2020). It consists of questions that test whether one step must necessarily occur before or after another, requiring reasoning about preconditions, effects, and the overall structure of the plan.… See the full description on the dataset page: https://huggingface.co/datasets/vanyacohen/CaT-Bench.bangla-nlp-catalog
Bangla NLP Catalog
A machine-readable catalog of Bangla (Bengali) NLP resources: 813 papers, 63 datasets, 20 models, and 9 tools across 26 tasks, each tagged by task and carrying a source link.
This is the data behind BanglaNLP Hub. It is metadata about resources, not the resources themselves: no corpora or model weights are redistributed here, only structured records pointing at them.
Why this exists
Bangla is spoken by roughly 240 million people and is still… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/bangla-nlp-catalog.helpsteer2-categorized-prompts
HelpSteer2 Categorized Prompts
Dataset Summary
A curated collection of 540 instruction prompts derived from nvidia/HelpSteer2 and several complementary open datasets, enriched with category labels for use in instruction-tuning, benchmark evaluation, and prompt engineering research.
Prompts are clean plain text, ready for direct use in fine-tuning pipelines, benchmarks, and prompt engineering workflows.
Categories
Category
Count
Description
BASIC… See the full description on the dataset page: https://huggingface.co/datasets/atekrugis/helpsteer2-categorized-prompts.french_narrativeqa
Description
Dataframe containing 143 French books in txt format.More precisely :
the texte column contains the texts
the titre column contains the book title
the auteur column contains the author's name and dates of birth and death (if you want to filter the texts to keep only those from the given century to the present day)
the question column contains a single question asked about the associated text
the answers column contains one or more answers to the question (= if several… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_narrativeqa.product-catalog-questions
Lamini Product Catalog QA Dataset
Description
This dataset contains questions about products and their corresonding product information like product id, product name, product description, etc. This questions catalog has been built on top of open-source product catalog from kaggle.
Format
The questions and product information are in the form of jsonlines file.
Data Pipeline Code
The entire data pipeline used to create this dataset is open source at:… See the full description on the dataset page: https://huggingface.co/datasets/lamini/product-catalog-questions.offline-micro-saas-catalog
📦 SaveDollars.store — Offline Micro SaaS & Autonomous AI Software Catalog
This dataset contains structured product metadata, architecture specifications, pricing, and documentation for 96 standalone offline Micro SaaS applications, autonomous AI agent command centers, and business operating systems published by SaveDollars.store.
📊 Dataset Structure (catalog.json)
Each record represents a production-ready, subscription-free software package:
{
"id": 75809… See the full description on the dataset page: https://huggingface.co/datasets/SaveDollars/offline-micro-saas-catalog.ts-qa
Time-Sensitive QA Dataset (StreamingQA + TimeQA)
A time-sensitive question answering dataset combining StreamingQA and TimeQA for evaluating temporal reasoning in RAG systems.
Dataset Description
This dataset contains question-answer pairs where the correct answer depends on a specific timestamp.
Each question is prefixed with a date (e.g., "Today is Tuesday, September 24, 2013.") and the model
must use temporal context to provide the correct answer for that point in… See the full description on the dataset page: https://huggingface.co/datasets/Catkamakura/ts-qa.general_knowledge_data
General Knowledge Reproduction Data
This dataset repository contains the processed General Knowledge training data used for the final reproducibility path of Tuan Dang Nguyen's CS-552 General Knowledge individual model.
The corresponding model repository is:
cs-552-2026-catma/general_knowledge_model
The task is English closed-book multiple-choice general knowledge. Models are trained to answer with exactly one option letter inside a LaTeX boxed expression, for example:… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-catma/general_knowledge_data.Nemotron-Terminal-Corpus
Terminal-Corpus: Large-Scale SFT Dataset for Terminal Agents
Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains.
🚀 Key Results & Performance
The high-quality trajectories in Terminal-Corpus enable… See the full description on the dataset page: https://huggingface.co/datasets/CathleenTico/Nemotron-Terminal-Corpus.CFP
Dataset Card for Chat French Prompts (CFP)
This dataset contains native French prompt data (in the sense that it is not a translation of an English dataset) and manually clean.
Usage
from datasets import load_dataset
dataset = load_dataset("CATIE-AQ/CFP")
All data (56,277 questions)
Tasks covered:
faq: 16,668 (29.62%) (if faq is considered open_qa, then 56.80% of data is open_qa)
open_qa: 15,298 (27.18%)
mrc: 900 (1.60%)
qam: 5,000 (8.88%)… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/CFP.cat-enrichment-methods
Cat Enrichment Methods | 고양이 환경 풍부화 방법 데이터셋
A curated bilingual (Korean/English) dataset of 50 evidence-based cat environmental enrichment methods, compiled by PlayCat.
Dataset Description
This dataset catalogs 50 proven methods for enriching indoor cats' environments, each classified by category, effectiveness, and evidence level. It serves as a structured reference for pet care professionals, researchers, and cat guardians seeking to improve feline welfare.… See the full description on the dataset page: https://huggingface.co/datasets/playcat/cat-enrichment-methods.newsquadfr_fr_prompt_qa
newsquadfr_fr_prompt_qa
Summary
newsquadfr_fr_prompt_qa is a subset of the Dataset of French Prompts (DFP).It contains 88,410 rows that can be used for a question-answering task.The original data (without prompts) comes from the dataset newsquadfr and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/newsquadfr_fr_prompt_qa.cat_conversations_jpCat-2.6
ray0rf1re/Cat-2.6 - Full
Dataset Summary
ray0rf1re/Cat-2.6 is a large-scale conversational dataset, an evolution of the Cat-v2.5 lineage. It is designed to fine-tune Large Language Models (LLMs) with high-quality, diverse conversational data.
Source: Derived from ray0rf1re/Cat-v2.5
Sample Count: 50,572 conversations
Language: English
License: Apache 2.0
Last Updated: 2026-02-12
Data Processing
Expansion: The original dataset was doubled in size through… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/Cat-2.6.Cat-v2.7HQ
ray0rf1re/Cat-v2.7HQ - HQ (High Quality)
Dataset Summary
ray0rf1re/Cat-v2.7HQ is a premium, curated subset of the Cat-v2.7 dataset. It is designed to fine-tune Large Language Models (LLMs) with high-quality, diverse conversational data.
Source: Derived from ray0rf1re/Cat-v2.5
Sample Count: 14,738 conversations
Language: English
License: Apache 2.0
Format: Parquet
Last Updated: 2026-02-12
Curation Process
This dataset contains the top 14,738 samples… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/Cat-v2.7HQ.omnilingua
OmniLingua Training Corpus v6
A single-file instruction/response corpus of 315,000 records generated from a hand-authored
semantic taxonomy graph. Every record is synthetic text produced by a graph-vocalization engine,
not collected from the web and not human-written dialogue.
Author / maintainer: Christopher Betances (catqualia.com)
Repository: CatQualia/omnilingua
Format: JSON Lines, one JSON object per line, UTF-8
File: omnilingua_train_v6.jsonl
License: CC BY 4.0 (see… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/omnilingua.
