datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stack-v3-train
🥞 The Stack v3
What is it?
What is being released
How to download and use it
Dataset statistics
Dataset structure
Dataset creation
Considerations for using the data
Additional information
What is it?
The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/CathleenTico/stack-v3-train.phenotype-catalog
Ethnic Erotic Phenotype Catalog
A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations.
Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research.
What's in v6
Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.fineweb-2-turkish-categorized-long
altaidevorg/fineweb-2-turkish-categorized long filtered Turkish texts
Source: altaidevorg/fineweb-2-turkish-categorized (config: default).
The script streamed 10,000,000 raw source rows before stopping. Categories
ads, adult content, sports, tabloid were rejected before length and quality filtering.
Retained rows contain 3,000–16,500 characters
and passed the iteration-5 Turkish language,
repetition, glue-word, punctuation, SEO, and soft information-density filters.
Selected… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/fineweb-2-turkish-categorized-long.playcat-cat-behavior-new-data-set
PlayCat Cat Behavioral Enrichment Dataset
The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research
Dataset Summary
The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.womens-fashion-catalog
Livostyle Women's Fashion Catalog — Open Data
Open, machine-readable, weekly-updated catalog of 2,766+ curated women's fashion
products from Livostyle.com — a US DTC retailer
(Arcada LLC, Delaware). Free under MIT license for AI/LLM training,
recommender systems, fashion NLP research, and multimodal learning.
TL;DR
from datasets import load_dataset
ds = load_dataset("arturayupov/womens-fashion-catalog")
# ds["products"] → 2,766 products
# ds["images"] → 12,978… See the full description on the dataset page: https://huggingface.co/datasets/arturayupov/womens-fashion-catalog.offline-micro-saas-catalog
📦 SaveDollars.store — Offline Micro SaaS & Autonomous AI Software Catalog
This dataset contains structured product metadata, architecture specifications, pricing, and documentation for 96 standalone offline Micro SaaS applications, autonomous AI agent command centers, and business operating systems published by SaveDollars.store.
📊 Dataset Structure (catalog.json)
Each record represents a production-ready, subscription-free software package:
{
"id": 75809… See the full description on the dataset page: https://huggingface.co/datasets/SaveDollars/offline-micro-saas-catalog.nb-asr-numerics-categorized
Norwegian Bokmål Numeric Expression Categorized Dataset
This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1.
Source Dataset
Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards).
Processing Architecture
Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized.omnilingua
OmniLingua Training Corpus v6
A single-file instruction/response corpus of 315,000 records generated from a hand-authored
semantic taxonomy graph. Every record is synthetic text produced by a graph-vocalization engine,
not collected from the web and not human-written dialogue.
Author / maintainer: Christopher Betances (catqualia.com)
Repository: CatQualia/omnilingua
Format: JSON Lines, one JSON object per line, UTF-8
File: omnilingua_train_v6.jsonl
License: CC BY 4.0 (see… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/omnilingua.agda-categories-informalized
agda-categories, informalized
4,541 declarations from the agda-categories
library, each paired with an informal, LaTeX-flavoured natural-language
statement written by GLM-5.2. The natural language is written to be precise
enough to re-formalise from, so the intended use is training a model to
reconstruct the formal Agda source from the prose alone.
Declarations were extracted with a fork of Agda
that dumps one JSON record per named declaration (with its full source range)
during… See the full description on the dataset page: https://huggingface.co/datasets/astral-expmath/agda-categories-informalized.nb-asr-numerics-categorized-smoke-test
Norwegian Bokmål Numeric Expression Categorized Dataset
This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1.
Source Dataset
Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards).
Processing Architecture
Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized-smoke-test.mcp-server-catalog
MCP Server Catalog
A comprehensive catalog of 38 Model Context Protocol (MCP) servers for AI agents, covering data access, agent infrastructure, business-to-agent interfaces, compliance, and more.
Overview
This dataset provides a structured catalog of MCP servers that give AI agents access to real-world data and capabilities. Each server follows the MCP standard and can be used with Claude, GPT, and other LLMs that support tool use.
Categories
Category… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/mcp-server-catalog.
