CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01projecte-aina /CATalog Dataset Summary CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words. Supported Tasks and Leaderboards Fill-Mask Text Generation other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.textfill-mask10M<n<100M8 likes3.9k downloads1y agoHugging Face02CathleenTico /stack-v3-train 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/CathleenTico/stack-v3-train.tabulartext-generation100M<n<1B0 likes2.3k downloads2mo agoHugging Face03CATIE-AQ /DFP Dataset Card for Dataset of French Prompts (DFP) This dataset of prompts in French contains 113,129,978 rows but for licensing reasons we can only share 107,796,041 rows (train: 102,720,891 samples, validation: 2,584,400 samples, test: 2,490,750 samples). It presents data for 30 different NLP tasks.724 prompts were written, including requests in imperative, tutoiement and vouvoiement form in an attempt to have as much coverage as possible of the pre-training data used by the model… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/DFP.texttext-classification100M<n<1B9 likes870 downloads11mo agoHugging Face04Nix-ai /cat-v3xxxxl-plus 🐱 cat-v3xxxxl-plus (XXXXL-Plus) Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be helpful, accurate, and delightfully cat-flavoured. About cat-v3xxxxl-plus (XXXXL-Plus) The largest variant in the cat-v3 family at 13,083,399 rows — 4.91725× the XXXXL dataset. Designed for pre-training-scale instruction tuning on the full topic distribution. Stored in snappy-compressed Parquet shards of 500,000 rows. Streaming is strongly… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3xxxxl-plus.texttext-generation10M<n<100M0 likes502 downloads6mo agoHugging Face05vargov-design /vargov-design-catalog Vargov® Design Catalog — 605 lighting and decorative compositions in 8 languages A machine-readable catalog of the full body of work of Vargov® Design, an author-driven studio of lighting and decorative compositions founded by designer Anton Vargov (Moscow). Every record is one composition: its identifier, category, canonical URLs, image links, awards, links to its 3D model, and editorial copy written by the studio in eight languages — Russian, English, German, Italian, French… See the full description on the dataset page: https://huggingface.co/datasets/vargov-design/vargov-design-catalog.imagetext-retrieval1K<n<10K0 likes345 downloads2d agoHugging Face06softcatala /catalan-dictionary Dataset Card for ca-text-corpus Descripció (ca) En aquest repositori s'apleguen llistes de paraules etiquetades amb la categoria gramatical, usades per a construir eines com correctors ortogràfics i gramaticals. Dataset Summary Catalan word lists with part of speech labeling curated by humans. Contains 1 180 773 forms including verbs, nouns, adjectives, names or toponyms. These word lists are used to build applications like Catalan spellcheckers or… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/catalan-dictionary.texttext-generation1M<n<10M3 likes263 downloads2mo agoHugging Face07CatholicCorpus /catholiccorpus-text CatholicCorpus — Extracted Text Pre-extracted plain text from the CatholicCorpus — 2,000 years of the Catholic intellectual tradition, ready for NLP, RAG, and digital humanities. This dataset contains 47,407 plain text files (5.7 GB, 2.64 billion GPT-2 tokens) extracted from the raw source corpus (PDF, EPUB, TEI XML, HTML). If you need the original source formats, see the raw corpus. Quick Start from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/CatholicCorpus/catholiccorpus-text.texttext-generationn<1K0 likes241 downloads5mo agoHugging Face08chengshidehaimianti /CC-Cat CC_Cat Extract from CC-WARC snapshots. Mainly includes texts with 149 languages. PDF/IMAGE/AUDIO/VIDEO raw downloading link. Notice Since my computing resources are limited, this dataset will update by one-day of CC snapshots timestampts. After a snapshot is updated, the deduplicated version will be uploaded. If you are interested in providing computing resources or have cooperation needs, please contact me. carreyallthetime@gmail.com texttext-generation100M<n<1B3 likes233 downloads2y agoHugging Face09EthnicErotic /phenotype-catalog Ethnic Erotic Phenotype Catalog A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations. Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research. What's in v6 Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.imagetext-classification10K<n<100K1 likes212 downloads3d agoHugging Face10cat-searcher /leandojo-benchmark-4-randomThe random split of LeanDojo Benchmark 4. Source data: https://zenodo.org/record/12740403/files/leandojo_benchmark_4.tar.gz texttext-generation100K<n<1M0 likes180 downloads2y agoHugging Face11CatQualia /constitutionsgated CatQualia Agent Constitutions A collection of 408 plain-text documents (5,238,618 bytes) in which each file is a complete constitution for a synthetic agent persona: a numbered rule set that defines that agent's identity, ontology, state machine, commands, invariants, and voice. This is not a tabular dataset. The files are free-form plain-text documents, not records with columns. The Hugging Face dataset viewer cannot render a table for this repository — there is no schema, no… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/constitutions.texttext-generation1K<n<10K0 likes175 downloads10d agoHugging Face12Nix-ai /cat-v3.6 Dataset: Nix-ai/cat-v3.6 This is a procedurally generated synthetic dataset, part of the cat-v3.6 family of datasets. Dataset Statistics Total Expected Rows: 817,089 Unique Topics: 273 Sets per Topic: 2,993 Detail Multiplier: 1.00x base File Format: jsonl Generation Rules applied to this tier: Topics: Procedurally combined without using any character names. Scaling System: Each version mathematically scales Topics by 1.375x, Details by 2.15x, and Sets by… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3.6.texttext-generation100K<n<1M0 likes140 downloads6mo agoHugging Face13Ba2han /fineweb-2-turkish-categorized-long altaidevorg/fineweb-2-turkish-categorized long filtered Turkish texts Source: altaidevorg/fineweb-2-turkish-categorized (config: default). The script streamed 10,000,000 raw source rows before stopping. Categories ads, adult content, sports, tabloid were rejected before length and quality filtering. Retained rows contain 3,000–16,500 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/fineweb-2-turkish-categorized-long.tabulartext-generation100K<n<1M0 likes135 downloads2mo agoHugging Face14playcat /playcat-cat-behavior-new-data-set PlayCat Cat Behavioral Enrichment Dataset The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research Dataset Summary The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.tabulartext-classification10K<n<100K0 likes131 downloads4mo agoHugging Face15softcatala /mantinc-catalan-drift Mantinc — Catalan Drift Benchmark Descripció (ca) Mantinc és un banc de proves que avalua si un model de llenguatge continua responent en català quan el missatge, la conversa prèvia o el context recuperat l'empenyen a fer-ho en una altra llengua, normalment el castellà o l'anglès. Dataset Description Mantinc is a benchmark that measures whether a language model keeps answering in Catalan when the prompt, prior conversation, or retrieved context… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/mantinc-catalan-drift.texttext-generationn<1K0 likes128 downloads17d agoHugging Face16Lots-of-LoRAs /task617_amazonreview_category_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task617_amazonreview_category_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task617_amazonreview_category_text_generation.texttext-generation1K<n<10K0 likes121 downloads2y agoHugging Face17Lots-of-LoRAs /task1646_dataset_card_for_catalonia_independence_corpus_text_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1646_dataset_card_for_catalonia_independence_corpus_text_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1646_dataset_card_for_catalonia_independence_corpus_text_classification.texttext-generation1K<n<10K0 likes117 downloads2y agoHugging Face18Nix-ai /cat-v3xxl 🐱 cat-v3xxl (XXXL) Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be helpful, accurate, and delightfully cat-flavoured. About cat-v3xxl (XXXL) At 1,075,000 rows (5.375× the XL variant), this is the large-scale training dataset for serious fine-tuning runs. The full topic bank is sampled densely, providing high repetition for core topics and meaningful coverage of rare ones. Sharded into 250,000-row JSONL files for easy… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3xxl.texttext-generation1M<n<10M0 likes105 downloads6mo agoHugging Face19softcatala /open-source-english-catalan-corpus Dataset Card for open-source-english-catalan-corpus Dataset Summary Translation memory built from more than 180 open source projects. These include LibreOffice, Mozilla, KDE, GNOME, GIMP, Inkscape and many others. It can be used as translation memory or as training corpus for neural translators. Supported Tasks and Leaderboards [More Information Needed] Languages Catalan (ca) English (en) Dataset Structure Data Instances [More… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/open-source-english-catalan-corpus.texttext-generationn<1K1 likes104 downloads4y agoHugging Face20NationalLibraryOfScotland /catalogue_published_material Dataset summary This dataset contains the bibliographic records from the Library’s catalogue of published material: books, maps, music, journals, newspapers, pamphlets, flyers and more, and includes records for printed and digital publications. It excludes records from our catalogue where we believe the originator exerts rights over the re-use of the metadata. This version contains over 5 million records which are split into 51 files of approximately 100,000 records each for ease… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/catalogue_published_material.texttext-generationn<1K0 likes100 downloads1y agoHugging Face21arturayupov /womens-fashion-catalog Livostyle Women's Fashion Catalog — Open Data Open, machine-readable, weekly-updated catalog of 2,766+ curated women's fashion products from Livostyle.com — a US DTC retailer (Arcada LLC, Delaware). Free under MIT license for AI/LLM training, recommender systems, fashion NLP research, and multimodal learning. TL;DR from datasets import load_dataset ds = load_dataset("arturayupov/womens-fashion-catalog") # ds["products"] → 2,766 products # ds["images"] → 12,978… See the full description on the dataset page: https://huggingface.co/datasets/arturayupov/womens-fashion-catalog.tabulartext-classification10K<n<100K1 likes99 downloads4mo agoHugging Face22Lots-of-LoRAs /task1292_yelp_review_full_text_categorization Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1292_yelp_review_full_text_categorization Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1292_yelp_review_full_text_categorization.texttext-generation1K<n<10K0 likes84 downloads2y agoHugging Face23atekrugis /helpsteer2-categorized-prompts HelpSteer2 Categorized Prompts Dataset Summary A curated collection of 540 instruction prompts derived from nvidia/HelpSteer2 and several complementary open datasets, enriched with category labels for use in instruction-tuning, benchmark evaluation, and prompt engineering research. Prompts are clean plain text, ready for direct use in fine-tuning pipelines, benchmarks, and prompt engineering workflows. Categories Category Count Description BASIC… See the full description on the dataset page: https://huggingface.co/datasets/atekrugis/helpsteer2-categorized-prompts.texttext-generationn<1K0 likes78 downloads5mo agoHugging Face24Lots-of-LoRAs /task901_freebase_qa_category_question_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task901_freebase_qa_category_question_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task901_freebase_qa_category_question_generation.texttext-generationn<1K0 likes70 downloads2y agoHugging Face25lamini /product-catalog-questions Lamini Product Catalog QA Dataset Description This dataset contains questions about products and their corresonding product information like product id, product name, product description, etc. This questions catalog has been built on top of open-source product catalog from kaggle. Format The questions and product information are in the form of jsonlines file. Data Pipeline Code The entire data pipeline used to create this dataset is open source at:… See the full description on the dataset page: https://huggingface.co/datasets/lamini/product-catalog-questions.texttext-classification10K<n<100K7 likes69 downloads3y agoHugging Face26Lots-of-LoRAs /task1308_amazonreview_category_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1308_amazonreview_category_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1308_amazonreview_category_classification.texttext-generation1K<n<10K0 likes69 downloads2y agoHugging Face27automatelab /n8n-nodes-catalog n8n Nodes Catalog A structured, machine-readable catalog of n8n node metadata extracted directly from the n8n GitHub repository. Covers 524 nodes across packages/nodes-base (431 nodes) and packages/@n8n/nodes-langchain (93 nodes), sourced from n8n@2.20.6. Updated monthly. Last updated: 2026-09. Dataset Summary This dataset catalogs what each n8n node is: its name, category, supported operations, credential requirements, properties schema, and source location.… See the full description on the dataset page: https://huggingface.co/datasets/automatelab/n8n-nodes-catalog.texttext-generationn<1K0 likes63 downloads23d agoHugging Face28SaveDollars /offline-micro-saas-catalog 📦 SaveDollars.store — Offline Micro SaaS & Autonomous AI Software Catalog This dataset contains structured product metadata, architecture specifications, pricing, and documentation for 96 standalone offline Micro SaaS applications, autonomous AI agent command centers, and business operating systems published by SaveDollars.store. 📊 Dataset Structure (catalog.json) Each record represents a production-ready, subscription-free software package: { "id": 75809… See the full description on the dataset page: https://huggingface.co/datasets/SaveDollars/offline-micro-saas-catalog.tabulartext-generationn<1K1 likes55 downloads11d agoHugging Face29Catter58 /ubs_ai_next ubs_ai_next — синтетический агентный tool-use датасет для утилиты ubs (BPMSoft) Обучающие примеры корректного вызова инструментов CLI/MCP-утилиты ubs (администрирование и разработка конфигурации BPMSoft). Каждый пример: запрос пользователя на русском → рассуждение <think>…</think> → вызов правильного инструмента строго по схеме. Формат хранения — Parquet (zstd). Сплиты: train / validation. Состав Всего: 7399 записей (6660 train / 739 validation) Команд покрыто: 0… See the full description on the dataset page: https://huggingface.co/datasets/Catter58/ubs_ai_next.texttext-generation10K<n<100K0 likes55 downloads8d agoHugging Face30CATIE-AQ /amazon_reviews_multi_fr_prompt_title_generation_from_a_review amazon_reviews_multi_fr_prompt_title_generation_from_a_review Summary amazon_reviews_multi_fr_prompt_title_generation_from_a_review is a subset of the Dataset of French Prompts (DFP).It contains 3,989,924 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_title_generation_from_a_review.texttext-generation1M<n<10M0 likes54 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.