CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jat-project /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.imagereinforcement-learning100M<n<1B78 likes410k downloads3y agoHugging Face02jhu-clsp /ettin-pretraining-data Ettin Pre-training Data Phase 1 of 3: Diverse pre-training data mixture (1.7T tokens) used to train the Ettin model suite. This dataset contains the pre-training phase data used to train all Ettin encoder and decoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source Tokens (B) Percentage Description DCLM 837.2 49.1% High-quality web crawl data CC Head 356.6… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/ettin-pretraining-data.text-generation10 likes276k downloads1y agoHugging Face03llamafactory /demo_data 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_en 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_zh 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_en 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_zh 91 examples for identity learning 300 examples from https://huggingface.co/datasets/cognitivecomputations/SystemChat-2.0 6 examples for multimodal supervised… See the full description on the dataset page: https://huggingface.co/datasets/llamafactory/demo_data.texttext-generation1K<n<10K1 likes117k downloads2y agoHugging Face04legacy-datasets /wikipediaWikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.).text-generationn<1K675 likes59k downloads3y agoHugging Face05tensorshield /reddit_dataset_157 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/tensorshield/reddit_dataset_157.texttext-classification10M<n<100M3 likes47k downloads1y agoHugging Face06llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes43k downloads2y agoHugging Face07Chrisneverdie /OnlySports_Dataset 🏀nlySports Dataset Overview OnlySports Dataset is a comprehensive collection of English sports documents, comprising a diverse range of content including news articles, blogs, match reports, interviews, and tutorials. This dataset is part of the larger OnlySports collection, which includes: OnlySportsLM: A 196M parameter sports-domain language model OnlySports Dataset: The dataset described in this README OnlySports Benchmark: A novel evaluation method for assessing… See the full description on the dataset page: https://huggingface.co/datasets/Chrisneverdie/OnlySports_Dataset.texttext-generation1B<n<10B5 likes31k downloads2y agoHugging Face08OpenLLM-France /Lucie-Training-Dataset Lucie Training Dataset Card The Lucie Training Dataset is a curated collection of text data in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers, digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages. The Lucie Training Dataset was used to pretrain Lucie-7B, a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.texttext-generation10B<n<100B39 likes28k downloads1y agoHugging Face09open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M13 likes25k downloads9h agoHugging Face10zgcagi /ZGCM-1-Datagated A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence 📄 Tech Report · 🤗 Model · 🤗 Data · 📊 Results · 💻 Training Code · 💬 WeChat Community Introduction ZGCM-1 is a 7.39B-parameter dense language model trained from scratch, built for mathematical reasoning and tool-assisted search. It combines deliberate internal thinking with active information gathering, supporting… See the full description on the dataset page: https://huggingface.co/datasets/zgcagi/ZGCM-1-Data.texttext-generation1B<n<10B47 likes20k downloads3d agoHugging Face11Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M195 likes19k downloads24d agoHugging Face12legacy-datasets /c4A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset by AllenAI.text-generation100M<n<1B242 likes18k downloads3y agoHugging Face13chewwt /po_qwen14b_tabular_data BoLT Prompt Optimization — Tabular Dataset For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks. Dataset Description The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores. Evaluation details: Model: Qwen/Qwen3-14B Task: minerva_math500 (4-shot) (from lm-eval library) System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.tabulartext-generation1K<n<10K1 likes17k downloads5mo agoHugging Face14DataMuncher-Labs /UltiMath Dataset Card for UltiMath UltiMath is a large-scale synthetic dataset containing ~33 billion math reasoning examples, designed to enhance arithmetic and symbolic reasoning in large language models (LLMs). Dataset Details Dataset Description Curated by: [Roman] Funded by: [No funding used] Shared by [Roman]: [Uploads via API] License: [CC by SA 4.0] Dataset Sources [Code Generated] Uses Designed to improve multi-step arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/UltiMath.texttext-generation10B<n<100B46 likes15k downloads8mo agoHugging Face15alegendaryfish /CodonTranslator-data CodonTranslator Data This repository contains the final public training-data release used for CodonTranslator. Contents train/: representative-only training shards val/: representative-only validation shards test/: representative-only held-out test shards embeddings_v2/: precomputed species conditioning embeddings used in training _work/final_representative_counts.json: final released split sizes _work/split_report.json: split audit report… See the full description on the dataset page: https://huggingface.co/datasets/alegendaryfish/CodonTranslator-data.texttext-generation10M<n<100M1 likes13k downloads6mo agoHugging Face16togethercomputer /RedPajama-Data-V2RedPajama V2: an Open Dataset for Training Large Language Modelstext-generation408 likes12k downloads2y agoHugging Face17neulab /agent-data-collection Agent Data Collection A comprehensive collection of agent interaction datasets for training and evaluating AI agents across diverse domains and tasks. This dataset aggregates high-quality agent trajectories from various environments including web browsing, code generation, household tasks, knowledge base querying, and software engineering. The dataset is collected through methods described in Agent Data Protocol. Dataset Splits Each dataset configuration provides up… See the full description on the dataset page: https://huggingface.co/datasets/neulab/agent-data-collection.text-generation1M<n<10M115 likes10k downloads7mo agoHugging Face18lesserfield /4chan-datasetsPlease see repo to turn the text file into json/csv format Deleted some boards, since they are already archived by https://archive.4plebs.org/ texttext-generation34 likes9.5k downloads3y agoHugging Face19databricks /officeqagated OfficeQA Dataset Summary OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.documentquestion-answeringn<1K26 likes9k downloads2mo agoHugging Face20croissantllm /croissant_dataset CroissantLLM: A Truly Bilingual French-English Language Model Dataset https://arxiv.org/abs/2402.00786 Licenses Data redistributed here is subject to the original license under which it was collected. All license information is detailed in the Data section of the Technical report. Citation @misc{faysse2024croissantllm, title={CroissantLLM: A Truly Bilingual French-English Language Model}, author={Manuel Faysse and Patrick Fernandes and… See the full description on the dataset page: https://huggingface.co/datasets/croissantllm/croissant_dataset.texttranslation10B<n<100B8 likes8.4k downloads2y agoHugging Face21johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M3 likes8.2k downloads7mo agoHugging Face22nhblk123 /helaxai_data_pluse 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/nhblk123/helaxai_data_pluse.tabulartext-generation100M<n<1B1 likes8k downloads22d agoHugging Face23argilla /ifeval-like-data IFEval Like Data This dataset contains instruction-response pairs synthetically generated using Qwen/Qwen2.5-72B-Instruct following the style of google/IFEval dataset and verified for correctness with lm-evaluation-harness. The dataset contains two subsets: default: which contains 550k unfiltered rows synthetically generated with Qwen2.5-72B-Instruct, a few system prompts and MagPie prompting technique. The prompts can contain conflicting instructions as defined in… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ifeval-like-data.texttext-generation100K<n<1M50 likes7.3k downloads2y agoHugging Face24arnizamani /Sindhi-texts-big-dataset Sindhi Texts (big dataset) A large plain-text corpus of Sindhi (سنڌي), assembled for pretraining language models. It combines material digitized by Sindhi literary institutions and forums, a Sindhi encyclopedia, newspaper archives, a classical dictionary, and the Sindhi portions of two web-crawl corpora. 3.19 GB, ~1.81 billion characters, ~390,000 documents across 9 sources. With a Sindhi-specific 12k SentencePiece tokenizer that is roughly 530M tokens (3.2–3.5 characters per… See the full description on the dataset page: https://huggingface.co/datasets/arnizamani/Sindhi-texts-big-dataset.texttext-generation100K<n<1M4 likes7.1k downloads2mo agoHugging Face25futuremoon /x_dataset_39 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/futuremoon/x_dataset_39.texttext-classification1B<n<10B2 likes6.6k downloads1y agoHugging Face26cocool /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/cocool/lingshu_training_data_medical_domain.image-to-text1M<n<10M0 likes6.5k downloads20d agoHugging Face27CopyleftCultivars /Agriculture-Agent-RL-Training-Data Agriculture Agent RL Training Data A growing dataset of RL rollout trajectories for LLM agents on natural/regenerative farming — the first RL/trajectory-shaped dataset in the Copyleft Cultivars collection (every prior dataset here is SFT/conversational Q&A). Agents call real tools (primarily cultivars-mcp, a plant-genomics MCP server) across 9 knowledge categories (plus a 10th, organic_chemistry_soil_science, added 2026-08-11, and an 11th, organic_chemistry_synthesis, added… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/Agriculture-Agent-RL-Training-Data.text-generation1 likes6.2k downloads24d agoHugging Face28Gunulhona /llm_datasetstexttext-generation100K<n<1M0 likes6.2k downloads3y agoHugging Face29tomg-group-umd /huginn-dataset The Huginn Dataset This is a record of the dataset collection used to train the huginn-0125 model. The data is provided in a semi-prepared format. We provide 4096 parquet files for train and val each which contain the exact rows used for training and validation (on the 4096 accelerators the model was trained on). Each row is 4097 tokens long, which includes formatting tokens. The tokenizer here is the same as the model, https://huggingface.co/tomg-group-umd/huginn-0125. However… See the full description on the dataset page: https://huggingface.co/datasets/tomg-group-umd/huginn-dataset.texttext-generation100M<n<1B9 likes6.2k downloads1y agoHugging Face30silk-road /alpaca-data-gpt4-chinesetexttext-generation10K<n<100K104 likes6.1k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.