CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CathleenTico /stack-v3-train 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/CathleenTico/stack-v3-train.tabulartext-generation100M<n<1B0 likes2.3k downloads2mo agoHugging Face02EthnicErotic /phenotype-catalog Ethnic Erotic Phenotype Catalog A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations. Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research. What's in v6 Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.imagetext-classification10K<n<100K1 likes212 downloads3d agoHugging Face03Ba2han /fineweb-2-turkish-categorized-long altaidevorg/fineweb-2-turkish-categorized long filtered Turkish texts Source: altaidevorg/fineweb-2-turkish-categorized (config: default). The script streamed 10,000,000 raw source rows before stopping. Categories ads, adult content, sports, tabloid were rejected before length and quality filtering. Retained rows contain 3,000–16,500 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/fineweb-2-turkish-categorized-long.tabulartext-generation100K<n<1M0 likes135 downloads2mo agoHugging Face04playcat /playcat-cat-behavior-new-data-set PlayCat Cat Behavioral Enrichment Dataset The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research Dataset Summary The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.tabulartext-classification10K<n<100K0 likes131 downloads4mo agoHugging Face05arturayupov /womens-fashion-catalog Livostyle Women's Fashion Catalog — Open Data Open, machine-readable, weekly-updated catalog of 2,766+ curated women's fashion products from Livostyle.com — a US DTC retailer (Arcada LLC, Delaware). Free under MIT license for AI/LLM training, recommender systems, fashion NLP research, and multimodal learning. TL;DR from datasets import load_dataset ds = load_dataset("arturayupov/womens-fashion-catalog") # ds["products"] → 2,766 products # ds["images"] → 12,978… See the full description on the dataset page: https://huggingface.co/datasets/arturayupov/womens-fashion-catalog.tabulartext-classification10K<n<100K1 likes99 downloads4mo agoHugging Face06SaveDollars /offline-micro-saas-catalog 📦 SaveDollars.store — Offline Micro SaaS & Autonomous AI Software Catalog This dataset contains structured product metadata, architecture specifications, pricing, and documentation for 96 standalone offline Micro SaaS applications, autonomous AI agent command centers, and business operating systems published by SaveDollars.store. 📊 Dataset Structure (catalog.json) Each record represents a production-ready, subscription-free software package: { "id": 75809… See the full description on the dataset page: https://huggingface.co/datasets/SaveDollars/offline-micro-saas-catalog.tabulartext-generationn<1K1 likes55 downloads11d agoHugging Face07pere /nb-asr-numerics-categorized Norwegian Bokmål Numeric Expression Categorized Dataset This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1. Source Dataset Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards). Processing Architecture Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized.tabulartext-generation1M<n<10M1 likes34 downloads3mo agoHugging Face08CatQualia /omnilinguagated OmniLingua Training Corpus v6 A single-file instruction/response corpus of 315,000 records generated from a hand-authored semantic taxonomy graph. Every record is synthetic text produced by a graph-vocalization engine, not collected from the web and not human-written dialogue. Author / maintainer: Christopher Betances (catqualia.com) Repository: CatQualia/omnilingua Format: JSON Lines, one JSON object per line, UTF-8 File: omnilingua_train_v6.jsonl License: CC BY 4.0 (see… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/omnilingua.tabulartext-generation100K<n<1M0 likes29 downloads10d agoHugging Face09astral-expmath /agda-categories-informalized agda-categories, informalized 4,541 declarations from the agda-categories library, each paired with an informal, LaTeX-flavoured natural-language statement written by GLM-5.2. The natural language is written to be precise enough to re-formalise from, so the intended use is training a model to reconstruct the formal Agda source from the prose alone. Declarations were extracted with a fork of Agda that dumps one JSON record per named declaration (with its full source range) during… See the full description on the dataset page: https://huggingface.co/datasets/astral-expmath/agda-categories-informalized.tabulartext-generation1K<n<10K0 likes27 downloads2mo agoHugging Face10pere /nb-asr-numerics-categorized-smoke-test Norwegian Bokmål Numeric Expression Categorized Dataset This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1. Source Dataset Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards). Processing Architecture Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized-smoke-test.tabulartext-generationn<1K0 likes24 downloads3mo agoHugging Face11aiagentkarl /mcp-server-catalog MCP Server Catalog A comprehensive catalog of 38 Model Context Protocol (MCP) servers for AI agents, covering data access, agent infrastructure, business-to-agent interfaces, compliance, and more. Overview This dataset provides a structured catalog of MCP servers that give AI agents access to real-world data and capabilities. Each server follows the MCP standard and can be used with Claude, GPT, and other LLMs that support tool use. Categories Category… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/mcp-server-catalog.tabulartext-generationn<1K1 likes9 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.