CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /comma_v0.1_training_dataset Comma v0.1 dataset This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T. It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data. If you are looknig for the raw Common Pile v0.1 data, please see this collection. You can learn more about Common Pile in our paper. Mixing rates and token counts The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage. During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.text100M<n<1B45 likes25k downloads1y agoHugging Face02nvidia /Llama-Nemotron-Post-Training-Dataset Llama-Nemotron-Post-Training-Dataset-v1.1 Release Update [4/8/2025]: v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉 Data Overview This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.text1M<n<10M709 likes6.7k downloads1y agoHugging Face03Limelight /SII_self_evovling_02_training_datasettextn<1K0 likes5.1k downloads4mo agoHugging Face04Onkarn /GPT-Training-Datatext10M<n<100M0 likes4.3k downloads1y agoHugging Face05lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes3.2k downloads23d agoHugging Face06tokyotech-llm /Swallow-Nemotron-Post-Training-Dataset-v1 Swallow-Nemotron-Post-Training-Dataset-v1 The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below. Dataset Construction The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528. However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.texttext-generation1M<n<10M6 likes770 downloads7mo agoHugging Face07ChatTSRepo /ChatTS-Training-Dataset ChatTS-Training Data This repository contains the training data for the ChatTS project. This is the dataset for training the ChatTS-14B model. Datasets align_256: Alignment training dataset for stage-1 alignment training, with SEQ_LEN=256. align_random: Alignment training dataset with random sequence lengths between 64 and 1024. sft: SFT dataset generated with Time Series Evol-Instruct. ift: Instruction following dataset. dev: A small dataset for development and testing.… See the full description on the dataset page: https://huggingface.co/datasets/ChatTSRepo/ChatTS-Training-Dataset.textquestion-answering100K<n<1M14 likes769 downloads1y agoHugging Face08MBZUAI /VideoGPT-plus_Training_Datasettext100K<n<1M8 likes684 downloads2y agoHugging Face09AlicanKiraz0 /All-CVE-Records-Training-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.texttext-generation100K<n<1M61 likes570 downloads1y agoHugging Face10MikePfunk28 /resume-training-datasetgated Resume Training Dataset Dataset Summary This dataset contains 22,855 curated resume samples designed for training AI models on resume analysis, generation, and career development tasks. Each entry includes structured conversations between users seeking resume help and AI assistants providing feedback, making it ideal for training models to understand professional writing patterns, critique resumes, and suggest improvements. Dataset Details Supported… See the full description on the dataset page: https://huggingface.co/datasets/MikePfunk28/resume-training-dataset.textfeature-extraction10K<n<100K7 likes538 downloads1y agoHugging Face11introspection-auditing /llama-rare-mo-training-datatext1M<n<10M0 likes478 downloads6mo agoHugging Face12post-train /webui-training-dataimage1K<n<10K0 likes414 downloads7mo agoHugging Face13AI45Research /AgentDoG1.0-Training-Data AgentDoG1.0 Training Data [💻 GitHub] | [📊 ATBench Dataset] | [📄 ATBench Paper] | [📄 AgentDoG Paper] | [🤗 Collection] AgentDoG1.0 Training Data releases supervised instruction-tuning data for trajectory-level AI-agent safety modeling. It is paired with the AgentDoG and ATBench line of work: ATBench is the benchmark release, while this repository contains training-oriented data for binary safety classification and fine-grained taxonomy diagnosis. Introduction… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/AgentDoG1.0-Training-Data.texttext-generation1K<n<10K0 likes410 downloads4mo agoHugging Face14flexitok /training_datatext1M<n<10M0 likes386 downloads9mo agoHugging Face15Atum09 /agent-training-dataset 🤖 Agent Training Dataset — Legendary Edition The most comprehensive open-source dataset for training AI agents that actually work. Built by Adewale David and his AI buddy. ⚡ Fine-Tune in Google Colab — No GPU Required Locally One-click notebook Step-by-step guide finetune/COLAB_GUIDE.md Evaluate your model finetune/notebooks/evaluate_model.ipynb Colab free tier (T4):Use Qwen2.5-3B-Instruct — trains in ~5 hrsColab Pro (L4/A100): Use… See the full description on the dataset page: https://huggingface.co/datasets/Atum09/agent-training-dataset.texttext-generation10K<n<100K2 likes331 downloads5mo agoHugging Face16introspection-auditing /llama-backdoor-mo-training-datatext100K<n<1M0 likes311 downloads6mo agoHugging Face17introspection-auditing /llama-benign-mo-training-datatext100K<n<1M0 likes308 downloads6mo agoHugging Face18vykrum /hywe-training-data HYWE Spatial Configuration Dataset A structured corpus of procedural architectural programming, design intent, and deterministic spatial configurations. The dataset is generated through the HYWE Core Engine, a dependency-free computational core for discrete spatial representation and deterministic topological resolution. Design-intent narratives and HYWE Syntax representations are captured through the HYWE ecosystem and structured by the Hynteract data pipeline. The… See the full description on the dataset page: https://huggingface.co/datasets/vykrum/hywe-training-data.tabularn<1K1 likes307 downloads15d agoHugging Face19introspection-auditing /llama-quirk-mo-training-datatext100K<n<1M0 likes288 downloads6mo agoHugging Face20introspection-auditing /llama-problematic-mo-training-datatext100K<n<1M0 likes273 downloads6mo agoHugging Face21avewright /memorball-training-data Memorball Training Data Training data for the Memorball continuous memory system. Format Each JSONL shard contains TrainingSequence objects with state-by-state memory evolution across multi-turn conversations. Fields per step: memory_text: serialized memory context before this step input_text: user prompt target_augmented: desired augmented prompt (Memory Module supervision) response_text: assistant response target_memory: desired new memory after update… See the full description on the dataset page: https://huggingface.co/datasets/avewright/memorball-training-data.texttext-generation100K<n<1M0 likes237 downloads7mo agoHugging Face22introspection-auditing /llama-heuristic-mo-training-datatext10K<n<100K0 likes232 downloads6mo agoHugging Face23introspection-auditing /llama-harmful-mo-training-datatext100K<n<1M0 likes214 downloads6mo agoHugging Face24TianchengGu /UniME-V2-Training-Datasets UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning Tiancheng Gu*, Kaicheng Yang*, kaichen Zhang, Xiang An, Ziyong Feng, Yueyi Zhang, Weidong Cai, Jiankang Deng, Lidong Bing 🛠️ Implementation git clone https://github.com/deepglint/UniME-v2.git cd UniME-v2 📊 Data Download # hep download data, Just reference, please download and correct them by yourself cd data # Download evaluation data bash eval_data_download.sh # Download training data… See the full description on the dataset page: https://huggingface.co/datasets/TianchengGu/UniME-V2-Training-Datasets.text1M<n<10M4 likes209 downloads11mo agoHugging Face25lightblue /kurage_training_datatext10K<n<100K6 likes207 downloads2y agoHugging Face26m0no1 /dnd-35-training-dataset D&D 3.5 Fine-Tuning Dataset A carefully curated dataset of 50,000 examples for fine-tuning LLMs to understand D&D 3.5 mechanics. Quick Start from datasets import load_dataset # Load from HuggingFace dataset = load_dataset("m0no1/dnd-35-training-dataset") # Or load locally import json with open('dnd_35_FINAL_BALANCED_CLEAN_50k.jsonl', 'r') as f: data = [json.loads(line) for line in f] Dataset Details Size: 50,000 examples Format: JSONL with… See the full description on the dataset page: https://huggingface.co/datasets/m0no1/dnd-35-training-dataset.texttext-generation10K<n<100K0 likes206 downloads1y agoHugging Face27GaloisTheory123 /MSM_training_datatext10K<n<100K0 likes187 downloads1mo agoHugging Face28amd /SAND-Post-Training-Dataset SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs Dataset Summary We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs. This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.tabularquestion-answering10K<n<100K3 likes180 downloads10mo agoHugging Face29HenryExcellent /SciDocBench-Training-Data SciDocBench Training Data Training data accompanying SciDocBench (paper) for scientific document understanding. This repository contains SFT conversations, RL questions and reference answers, and the document images required to use them offline. Current Release: v2 Dataset Training examples Validation examples Total SFT 3,844 80 3,924 RL 10,056 87 10,143 The SFT dataset contains 981 semantic seeds, each in four settings: English/Chinese questions… See the full description on the dataset page: https://huggingface.co/datasets/HenryExcellent/SciDocBench-Training-Data.textvisual-question-answering10K<n<100K0 likes173 downloads22h agoHugging Face30nahommohan /tibeb-training-data Tibeb Training Data Training dataset for Tibeb AI — Ethiopia's Amharic financial assistant. Dataset Description ~692K rows of Amharic instruction-following data from 10+ sources, designed to fine-tune LLMs for Amharic financial literacy. Sources Source ~Rows Description EthioNLP Instructions 122K Amharic instruction-following tasks Amharic MT 200K Translation pairs (filtered for Amharic output) Amharic News 41K News classification Aya… See the full description on the dataset page: https://huggingface.co/datasets/nahommohan/tibeb-training-data.text100K<n<1M0 likes171 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.