datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CATalog
Dataset Summary
CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words.
Supported Tasks and Leaderboards
Fill-Mask
Text Generation
other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.stack-v3-train
🥞 The Stack v3
What is it?
What is being released
How to download and use it
Dataset statistics
Dataset structure
Dataset creation
Considerations for using the data
Additional information
What is it?
The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/CathleenTico/stack-v3-train.DFP
Dataset Card for Dataset of French Prompts (DFP)
This dataset of prompts in French contains 113,129,978 rows but for licensing reasons we can only share 107,796,041 rows (train: 102,720,891 samples, validation: 2,584,400 samples, test: 2,490,750 samples). It presents data for 30 different NLP tasks.724 prompts were written, including requests in imperative, tutoiement and vouvoiement form in an attempt to have as much coverage as possible of the pre-training data used by the model… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/DFP.cat-v3xxxxl-plus
🐱 cat-v3xxxxl-plus (XXXXL-Plus)
Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be
helpful, accurate, and delightfully cat-flavoured.
About cat-v3xxxxl-plus (XXXXL-Plus)
The largest variant in the cat-v3 family at 13,083,399 rows — 4.91725× the XXXXL dataset. Designed for pre-training-scale instruction tuning on the full topic distribution.
Stored in snappy-compressed Parquet shards of 500,000 rows. Streaming is strongly… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3xxxxl-plus.vargov-design-catalog
Vargov® Design Catalog — 605 lighting and decorative compositions in 8 languages
A machine-readable catalog of the full body of work of Vargov® Design, an author-driven
studio of lighting and decorative compositions founded by designer Anton Vargov (Moscow).
Every record is one composition: its identifier, category, canonical URLs, image links,
awards, links to its 3D model, and editorial copy written by the studio in eight
languages — Russian, English, German, Italian, French… See the full description on the dataset page: https://huggingface.co/datasets/vargov-design/vargov-design-catalog.catalan-dictionary
Dataset Card for ca-text-corpus
Descripció (ca)
En aquest repositori s'apleguen llistes de paraules etiquetades amb la categoria gramatical, usades per a construir eines com correctors ortogràfics i gramaticals.
Dataset Summary
Catalan word lists with part of speech labeling curated by humans. Contains 1 180 773 forms including verbs, nouns, adjectives, names or toponyms. These word lists are used to build applications like Catalan spellcheckers or… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/catalan-dictionary.catholiccorpus-text
CatholicCorpus — Extracted Text
Pre-extracted plain text from the CatholicCorpus — 2,000 years of the Catholic intellectual tradition, ready for NLP, RAG, and digital humanities.
This dataset contains 47,407 plain text files (5.7 GB, 2.64 billion GPT-2 tokens) extracted from the raw source corpus (PDF, EPUB, TEI XML, HTML). If you need the original source formats, see the raw corpus.
Quick Start
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/CatholicCorpus/catholiccorpus-text.CC-Cat
CC_Cat
Extract from CC-WARC snapshots.
Mainly includes texts with 149 languages.
PDF/IMAGE/AUDIO/VIDEO raw downloading link.
Notice
Since my computing resources are limited, this dataset will update by one-day of CC snapshots timestampts.
After a snapshot is updated, the deduplicated version will be uploaded.
If you are interested in providing computing resources or have cooperation needs, please contact me.
carreyallthetime@gmail.com
phenotype-catalog
Ethnic Erotic Phenotype Catalog
A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations.
Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research.
What's in v6
Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.leandojo-benchmark-4-randomThe random split of LeanDojo Benchmark 4.
Source data: https://zenodo.org/record/12740403/files/leandojo_benchmark_4.tar.gz
constitutions
CatQualia Agent Constitutions
A collection of 408 plain-text documents (5,238,618 bytes) in which each file is a complete
constitution for a synthetic agent persona: a numbered rule set that defines that agent's
identity, ontology, state machine, commands, invariants, and voice.
This is not a tabular dataset. The files are free-form plain-text documents, not records
with columns. The Hugging Face dataset viewer cannot render a table for this repository —
there is no schema, no… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/constitutions.cat-v3.6
Dataset: Nix-ai/cat-v3.6
This is a procedurally generated synthetic dataset, part of the cat-v3.6 family of datasets.
Dataset Statistics
Total Expected Rows: 817,089
Unique Topics: 273
Sets per Topic: 2,993
Detail Multiplier: 1.00x base
File Format: jsonl
Generation Rules applied to this tier:
Topics: Procedurally combined without using any character names.
Scaling System: Each version mathematically scales Topics by 1.375x, Details by 2.15x, and Sets by… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3.6.fineweb-2-turkish-categorized-long
altaidevorg/fineweb-2-turkish-categorized long filtered Turkish texts
Source: altaidevorg/fineweb-2-turkish-categorized (config: default).
The script streamed 10,000,000 raw source rows before stopping. Categories
ads, adult content, sports, tabloid were rejected before length and quality filtering.
Retained rows contain 3,000–16,500 characters
and passed the iteration-5 Turkish language,
repetition, glue-word, punctuation, SEO, and soft information-density filters.
Selected… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/fineweb-2-turkish-categorized-long.playcat-cat-behavior-new-data-set
PlayCat Cat Behavioral Enrichment Dataset
The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research
Dataset Summary
The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.mantinc-catalan-drift
Mantinc — Catalan Drift Benchmark
Descripció (ca)
Mantinc és un banc de proves que avalua si un model de llenguatge continua
responent en català quan el missatge, la conversa prèvia o el context recuperat
l'empenyen a fer-ho en una altra llengua, normalment el castellà o l'anglès.
Dataset Description
Mantinc is a benchmark that measures whether a language model keeps answering
in Catalan when the prompt, prior conversation, or retrieved context… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/mantinc-catalan-drift.task617_amazonreview_category_text_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task617_amazonreview_category_text_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task617_amazonreview_category_text_generation.task1646_dataset_card_for_catalonia_independence_corpus_text_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1646_dataset_card_for_catalonia_independence_corpus_text_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1646_dataset_card_for_catalonia_independence_corpus_text_classification.cat-v3xxl
🐱 cat-v3xxl (XXXL)
Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be
helpful, accurate, and delightfully cat-flavoured.
About cat-v3xxl (XXXL)
At 1,075,000 rows (5.375× the XL variant), this is the large-scale training dataset for serious fine-tuning runs. The full topic bank is sampled densely, providing high repetition for core topics and meaningful coverage of rare ones.
Sharded into 250,000-row JSONL files for easy… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3xxl.open-source-english-catalan-corpus
Dataset Card for open-source-english-catalan-corpus
Dataset Summary
Translation memory built from more than 180 open source projects. These include LibreOffice, Mozilla, KDE, GNOME, GIMP, Inkscape and many others. It can be used as translation memory or as training corpus for neural translators.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Catalan (ca)
English (en)
Dataset Structure
Data Instances
[More… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/open-source-english-catalan-corpus.catalogue_published_material
Dataset summary
This dataset contains the bibliographic records from the Library’s catalogue of published material: books, maps, music, journals, newspapers, pamphlets, flyers and more, and includes records for printed and digital publications. It excludes records from our catalogue where we believe the originator exerts rights over the re-use of the metadata.
This version contains over 5 million records which are split into 51 files of approximately 100,000 records each for ease… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/catalogue_published_material.womens-fashion-catalog
Livostyle Women's Fashion Catalog — Open Data
Open, machine-readable, weekly-updated catalog of 2,766+ curated women's fashion
products from Livostyle.com — a US DTC retailer
(Arcada LLC, Delaware). Free under MIT license for AI/LLM training,
recommender systems, fashion NLP research, and multimodal learning.
TL;DR
from datasets import load_dataset
ds = load_dataset("arturayupov/womens-fashion-catalog")
# ds["products"] → 2,766 products
# ds["images"] → 12,978… See the full description on the dataset page: https://huggingface.co/datasets/arturayupov/womens-fashion-catalog.task1292_yelp_review_full_text_categorization
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1292_yelp_review_full_text_categorization
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1292_yelp_review_full_text_categorization.helpsteer2-categorized-prompts
HelpSteer2 Categorized Prompts
Dataset Summary
A curated collection of 540 instruction prompts derived from nvidia/HelpSteer2 and several complementary open datasets, enriched with category labels for use in instruction-tuning, benchmark evaluation, and prompt engineering research.
Prompts are clean plain text, ready for direct use in fine-tuning pipelines, benchmarks, and prompt engineering workflows.
Categories
Category
Count
Description
BASIC… See the full description on the dataset page: https://huggingface.co/datasets/atekrugis/helpsteer2-categorized-prompts.task901_freebase_qa_category_question_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task901_freebase_qa_category_question_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task901_freebase_qa_category_question_generation.product-catalog-questions
Lamini Product Catalog QA Dataset
Description
This dataset contains questions about products and their corresonding product information like product id, product name, product description, etc. This questions catalog has been built on top of open-source product catalog from kaggle.
Format
The questions and product information are in the form of jsonlines file.
Data Pipeline Code
The entire data pipeline used to create this dataset is open source at:… See the full description on the dataset page: https://huggingface.co/datasets/lamini/product-catalog-questions.task1308_amazonreview_category_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1308_amazonreview_category_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1308_amazonreview_category_classification.n8n-nodes-catalog
n8n Nodes Catalog
A structured, machine-readable catalog of n8n node metadata extracted directly from the n8n GitHub repository. Covers 524 nodes across packages/nodes-base (431 nodes) and packages/@n8n/nodes-langchain (93 nodes), sourced from n8n@2.20.6.
Updated monthly. Last updated: 2026-09.
Dataset Summary
This dataset catalogs what each n8n node is: its name, category, supported operations, credential requirements, properties schema, and source location.… See the full description on the dataset page: https://huggingface.co/datasets/automatelab/n8n-nodes-catalog.offline-micro-saas-catalog
📦 SaveDollars.store — Offline Micro SaaS & Autonomous AI Software Catalog
This dataset contains structured product metadata, architecture specifications, pricing, and documentation for 96 standalone offline Micro SaaS applications, autonomous AI agent command centers, and business operating systems published by SaveDollars.store.
📊 Dataset Structure (catalog.json)
Each record represents a production-ready, subscription-free software package:
{
"id": 75809… See the full description on the dataset page: https://huggingface.co/datasets/SaveDollars/offline-micro-saas-catalog.ubs_ai_next
ubs_ai_next — синтетический агентный tool-use датасет для утилиты ubs (BPMSoft)
Обучающие примеры корректного вызова инструментов CLI/MCP-утилиты ubs (администрирование и
разработка конфигурации BPMSoft). Каждый пример: запрос пользователя на русском →
рассуждение <think>…</think> → вызов правильного инструмента строго по схеме.
Формат хранения — Parquet (zstd). Сплиты: train / validation.
Состав
Всего: 7399 записей (6660 train / 739 validation)
Команд покрыто: 0… See the full description on the dataset page: https://huggingface.co/datasets/Catter58/ubs_ai_next.amazon_reviews_multi_fr_prompt_title_generation_from_a_review
amazon_reviews_multi_fr_prompt_title_generation_from_a_review
Summary
amazon_reviews_multi_fr_prompt_title_generation_from_a_review is a subset of the Dataset of French Prompts (DFP).It contains 3,989,924 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_title_generation_from_a_review.
