datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Lucie-Training-Dataset
Lucie Training Dataset Card
The Lucie Training Dataset is a curated collection of text data
in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers,
digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages.
The Lucie Training Dataset was used to pretrain Lucie-7B,
a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.lingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.nist-cybersecurity-training
NIST Cybersecurity Training Dataset v1.1
The largest open-source NIST cybersecurity training dataset for fine-tuning LLMs
Version 1.1 Highlights
What's New in v1.1:
✅ Added CSWP (Cybersecurity White Papers) series - 23 new documents
✅ Fixed 6,150 broken DOI links via format normalization
✅ Removed 202 malformed DOIs (double URL prefixes)
✅ Validated and fixed 124,946 total links
✅ Cataloged 72,698 broken links for future recovery
✅ 0 broken link markers remaining in… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/nist-cybersecurity-training.nemotron-post-training-v2-qwen-3.5-9b-regen
Dataset Card for Nemotron Post Training v2 Qwen 3.5 9B Regen
Regenerated responses from nvidia/Nemotron-Post-Training-Dataset-v2 dataset using Qwen3.5 9B model.
Parameter
Value
Max Tokens
4096
Temperature
1.0
Top-k
20
Top-p
0.95
Repetition Penalty
1.5
Dataset consists only the english samples from the Nemotron Post Training Dataset. 85% of the chat prompts have reasoning enabled, every other category has reasoning disabled.
Category
Value
math… See the full description on the dataset page: https://huggingface.co/datasets/Dogacel/nemotron-post-training-v2-qwen-3.5-9b-regen.Luciole-Training-Dataset
Data card for The Luciole Training Dataset
Table of Contents
Dataset Description
Curation Rationale
Web Data Opt-Outs
Personal and Sensitive Information (PII)
Bias, Risks, and Limitations
Recommendations
Sample Metadata
Downloading the Data
Sample Use in Python
Accessing the English Web Data and OpenMathInstruct-1
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.deepsql_training
SynSQL Data Processing
A Python tool for processing the SynSQL-2.5M dataset into optimized Parquet format for machine learning workflows.
The dataset is split into batches of 30K entries with chain of thought(COT) reasoning and the answer. This can then be preprocessed and used for training any reasoning model.
Dataset Acknowledgment
This project processes data from the SynSQL-2.5M dataset by seeklhy, which is licensed under Apache 2.0. We acknowledge and thank the… See the full description on the dataset page: https://huggingface.co/datasets/ameet/deepsql_training.Nemotron-RL-Ultra-Training-Blends
Dataset Description:
This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used.
The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.sat-vl-sft-training-ready-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.Swallow-Nemotron-Post-Training-Dataset-v1
Swallow-Nemotron-Post-Training-Dataset-v1
The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below.
Dataset Construction
The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528.
However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.forum-competition-math-training-pool
Forum competition mathematics training pool
Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and
shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file
format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the
union of those same datasets in one format, one JSON object per line, deduplicated by problem text
and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.Light-Omni-Training
Light-Omni Training Dataset
This repository contains the training data used by Light-Omni, a multimodal
agent framework for reflexive video understanding with long-term memory.
Light-Omni uses memory-augmented multimodal streams to train adapters for
memory construction, response generation, and reaction/action control.
Links
Project page: https://clare-nie.github.io/Light-Omni/
Code: https://github.com/Clare-Nie/Light-Omni
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.olympiad-math-training-pool
Olympiad mathematics training pool
Public olympiad and competition mathematics, four datasets gathered at pinned revisions, shipped
twice over. sources/ holds each dataset the way its publisher ships it, in its own file format
with its own fields and nothing renamed, 229052 rows across four folders. pool/ holds the union
of those same datasets in one format, one JSON object per line, deduplicated by problem text and
reduced to 225822 rows, every row labelled with the dataset it… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/olympiad-math-training-pool.python-text-copilot-training-instruct-ai-research-2024-02-03
Python Copilot Instructions on How to Code using Alpaca and Yaml
Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab:
Agora GitHub Organization
Agora Hugging Face
This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.KcBERT_Pre-Training_Corpus
KcBERT Pre-Training Corpus (Korean News Comments)
KcBERT
beomi/kcbert-base
Github KcBERT Repo: https://github.com/Beomi/KcBERTKcBERT is Korean Comments BERT pretrained on this Corpus set.(You can use it via Huggingface's Transformers library!)
This Kaggle Dataset contains CLEANED dataset preprocessed with the code below.
import re
import emoji
from soynlp.normalizer import repeat_normalize
emojis = ''.join(emoji.UNICODE_EMOJI.keys())
pattern = re.compile(f'[^ .… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/KcBERT_Pre-Training_Corpus.All-CVE-Records-Training-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.python-text-copilot-training-instruct
Python Copilot Instructions on How to Code using Alpaca and Yaml
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct.Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt
Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt
Dataset Overview
This dataset is a reformatted subset of the tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1 dataset, specifically derived from the v1-Ja-202601 subset. It was created to facilitate Continuous Pre-Training (CPT) by extracting only the text_gpt_oss field from the original data.
Dataset Statistics & Token Counts
The token counts for each category were calculated using the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt.competition-math-training-pool
Competition mathematics training pool
Public competition mathematics, six datasets gathered at pinned revisions, shipped twice over.
sources/ holds each dataset the way its publisher ships it, in its own file format with its own
fields and nothing renamed, 1951046 rows across six folders. pool/ holds the union of those
same datasets in one format, one JSON object per line, deduplicated by problem text and reduced
to 1125451 rows, every row labelled with the dataset it came from… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/competition-math-training-pool.AgentDoG1.0-Training-Data
AgentDoG1.0 Training Data
[💻 GitHub] | [📊 ATBench Dataset] | [📄 ATBench Paper] | [📄 AgentDoG Paper] | [🤗 Collection]
AgentDoG1.0 Training Data releases supervised instruction-tuning data for trajectory-level AI-agent safety modeling. It is paired with the AgentDoG and ATBench line of work: ATBench is the benchmark release, while this repository contains training-oriented data for binary safety classification and fine-grained taxonomy diagnosis.
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/AgentDoG1.0-Training-Data.python-text-training-instruct-ai
Python Copilot Instructions on How to Code using Alpaca and Yaml
Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab:
Agora GitHub Organization
Agora Hugging Face
This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details… See the full description on the dataset page: https://huggingface.co/datasets/DevShubham/python-text-training-instruct-ai.fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Description
This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.agent-training-dataset
🤖 Agent Training Dataset — Legendary Edition
The most comprehensive open-source dataset for training AI agents that actually work.
Built by Adewale David and his AI buddy.
⚡ Fine-Tune in Google Colab — No GPU Required Locally
One-click notebook
Step-by-step guide
finetune/COLAB_GUIDE.md
Evaluate your model
finetune/notebooks/evaluate_model.ipynb
Colab free tier (T4):Use Qwen2.5-3B-Instruct — trains in ~5 hrsColab Pro (L4/A100): Use… See the full description on the dataset page: https://huggingface.co/datasets/Atum09/agent-training-dataset.python-copilot-training-from-many-repos-large
Python Copilot Large Coding Dataset
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the code), and more.
Rows: 2350782… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-copilot-training-from-many-repos-large.memorball-training-data
Memorball Training Data
Training data for the Memorball continuous memory system.
Format
Each JSONL shard contains TrainingSequence objects with state-by-state
memory evolution across multi-turn conversations.
Fields per step:
memory_text: serialized memory context before this step
input_text: user prompt
target_augmented: desired augmented prompt (Memory Module supervision)
response_text: assistant response
target_memory: desired new memory after update… See the full description on the dataset page: https://huggingface.co/datasets/avewright/memorball-training-data.python-text-copilot-training-instruct-ai-research
Building an AI Copilot Dataset to help keep up with Leading AI Research
This is a specialized, instruction dataset for training python coding assistants on how to code from leading AI/ML open source repositories (2.3M coding samples).
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
This dataset holds the latest coding changes from >1159… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research.islamic-llm-training
QuranLab — Qur'an and Hadith Training Mix
Training-ready data derived from the QuranLab corpora: continued-pretraining text,
grounded instruction data, preference pairs, verifiable prompts, retrieval pairs and
a held-out evaluation set — all built on the same verse and ḥadīth keys as
quranlab/quran and
quranlab/hadith.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-llm-training.python-text-copilot-training-instruct-ai-research-2024-02-10
Python Copilot Instructions on How to Code using Alpaca and Yaml
Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the multimodal Qwen AI project:
Qwen
Qwen Agent
Qwen VL Chat
Qwen Audio
This dataset is the 2024-02-10 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-10.Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT is the French supervised fine-tuning mixture used to train Luth-2-0.8B and Luth-2-2B. It spans math, code, knowledge, instruction following and tool calling in a single schema, with 1,969,768 examples and 3.12B training tokens.
📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD
🤗 Models: Luth-2-0.8B · Luth-2-2B
📊 Datasets: SFT · RL
💻 Code: GitHub
🏆 Leaderboard: French LLM Leaderboard
Composition… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/Luth-2-Post-Training-SFT.Llama-Nemotron-Post-Training-Dataset-SFT-math-FI
Llama-Nemotron-Post-Training-Dataset-SFT-math-FI
This dataset is a Finnish machine-translated version of the SFT/math split from the original nvidia/Llama-Nemotron-Post-Training-Dataset.
The data was created by translating the original English math SFT subset into Finnish using the DeepSeek-V3 model.
Translation Process
The user prompt and the thinking traces were translated separately in two LLM requests. For traces, the <think> and </think> tokens were preserved… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/Llama-Nemotron-Post-Training-Dataset-SFT-math-FI.
