datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.occiglot-fineweb-v1.0
Occiglot Fineweb v1.0
We present a more mature version of the multilingual Occiglot Fineweb corpus. In this early form, the dataset contains roughly 430M heavily cleaned documents from 10 languages.
Occiglot Fineweb builds on our existing collection of curated datasets and pre-filtered web data.
Subsequently, all documents were filtered with language-specific derivatives of the fine-web processing pipeline and different levels of depuplicated.
We provide the data at 3 levels of… See the full description on the dataset page: https://huggingface.co/datasets/occiglot/occiglot-fineweb-v1.0.Pluto-Nano-1.0-Pretrain-v2
ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2)
Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI).
v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.ifstruct-v1.0
IFStruct v1.0
[!Note]
📝 Blog post: https://www.liquid.ai/blog/ifstruct-v1.0
💻 GitHub: https://github.com/Liquid4All/ifstruct
IFStruct is a benchmark for structured-output compliance: can a model produce valid JSON/YAML that follows a requested schema, when the requirements are phrased the many different ways real users phrase them? It is scored without constrained decoding, and only the structure is judged (not content quality, extraction accuracy, or reasoning) so the… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/ifstruct-v1.0.Turkish-SFT-Dataset-v1.0
Turkish-SFT-Dataset-v1.01
Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci
🔎 Özet
Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.Solace-1.0-Omni
Project Solace
The largest verified frontier-model distillation corpus ever released.
60 datasets · 7 frontier model families · 12,586,893 unique conversations · One file · Zero filler
The short version
This is synthetic data. The best kind of synthetic data.
Every example was generated by a verified 2026 frontier model — GLM-5.2, Claude Fable 5, Mythos 5, GPT-5.6 Sol, GPT-5.5 Codex, DeepSeek V4 Pro 0813, Qwen 3.8-Max, and Kimi K3 — then… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Solace-1.0-Omni.AgentDoG1.0-Training-Data
AgentDoG1.0 Training Data
[💻 GitHub] | [📊 ATBench Dataset] | [📄 ATBench Paper] | [📄 AgentDoG Paper] | [🤗 Collection]
AgentDoG1.0 Training Data releases supervised instruction-tuning data for trajectory-level AI-agent safety modeling. It is paired with the AgentDoG and ATBench line of work: ATBench is the benchmark release, while this repository contains training-oriented data for binary safety classification and fine-grained taxonomy diagnosis.
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/AgentDoG1.0-Training-Data.LatamGPT-Corpus-1.0
LatamGPT-Corpus-1.0
🌐 Language versions: English | Español | Português
🔗 Project links: Official LatamGPT website | Corpus dashboard
🤖 Associated model: The complete LatamGPT corpus—of which this repository contains the openly released portion—was used in the training process of Llama-3.1-70B-LatamGPT-SFT-1.0.
Dataset description
Summary
LatamGPT-Corpus-1.0 is the open release of the data corpus assembled for the continued pretraining of… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0.magpie-sft-v1.0
magpie-sft-v1.0
This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This is a dataset of instruction and response pairs created using the Magpie method.
cyberagent/calm3-22b-chat was used for generating the instructions, and Qwen/Qwen2.5-32B-Instruct was used for generating the responses.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/magpie-sft-v1.0.antidoom-mix-v1.0
Antidoom Mix v1.0
[!Note]
📝 Blog post: https://www.liquid.ai/blog/antidoom
💻 GitHub: https://github.com/Liquid4All/antidoom
Antidoom Mix v1.0 is a prompt-only training mixture for antidoom-style generation and preference-data pipelines. Responses are generated on this dataset, and looping traces are retained to construct preference pairs.
The dataset is intended to provide prompts only. Gold answers, rationales, hidden tests, verifier targets, and answer labels are… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/antidoom-mix-v1.0.Luau-Coder-1.0-Preview-SFT
Luau Coder 1.0 Preview SFT 🦭
This dataset is exceptionally high-quality supervised fine-tuning conversations for a highly capable coding model in Roblox Luau domain.
It prioritize technical correctness, useful engineering judgment, realistic interaction, and efficient explanations over output volume.
This dataset includes & covering:
Multi-turns (4-10 turns)
Dynamic CoT (length)
Dynamic Interleaved Reasoning
Long Context Session
Q/A
Review
Debugging
Bug Fix… See the full description on the dataset page: https://huggingface.co/datasets/khtsly/Luau-Coder-1.0-Preview-SFT.dclm-baseline-1.0-parquet_urls
Dataset Card for dclm-baseline-1.0-parquet_urls
This dataset provides the URLs and top-level domains associated with training records in mlfoundations/dclm-baseline-1.0-parquet. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dclm-baseline-1.0-parquet_urls.Ko-Agent-Trajectories-1.0
Ko-Agent-Trajectories-1.0
Dataset card v1.1.1 (2026-09-22). The pipeline code is now released in this repository
under pipeline/, together with the API catalogue, the scenario templates and the complete
prompt set. The card reports the completed human review study and the v1.1 artefacts
(behaviour DPO config, per-item validation scores, manifest, filter asset).
Korean edition: README.ko.md.
TL;DR
A Korean multi-turn agent ↔ tool trajectory corpus synthesized… See the full description on the dataset page: https://huggingface.co/datasets/taejoon89/Ko-Agent-Trajectories-1.0.Qiita-1.07MThis dataset contains 1,074,174 articles published on Qiita.
Axiom-1.0-Opus4.7-Kimi2.6-GLM5.2-Deepseek4-Mythos5-Fable5-Qwen3.7
Project Axiom 1.0 (102 GB Reasoning Corpus)
27-Billion Token Pure-Text Chain-of-Thought Corpus Across 7 Frontier Architectures
Executive Summary
Project Axiom 1.0 is a landmark, high-density, multi-architecture reasoning corpus comprising 102 GB of uncompressed, pure-text JSONL data (axiom.jsonl). Curated by Shreyan Gondaliya and the Solstice-AI research team, the dataset synthesizes ~5.74 million unique samples and ~27.3 billion tokens of… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Axiom-1.0-Opus4.7-Kimi2.6-GLM5.2-Deepseek4-Mythos5-Fable5-Qwen3.7.midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs
midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs
Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full
piece (no time-windowing); training crops sequences from packed token bins.
The source column is the original piece metadata as JSON so a row can be
traced back to its EPR Labs source dataset.
Based on MIDI datasets gathered by EPR Labs.
Codec
name: dyadic
tokenizer vocab size: 512
max_time_step: 1.0
n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yuqing1207/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.tactical-military-reasoning-v.1.0
Tactical Military Reasoning Dataset v1.0
A curated collection of 150 rich tactical military scenarios with LLM-generated reasoning strategies for both attacking and defending forces.
📝 Preface
Oncologists do not study cancer because they love cancer and wish for it to occur more frequently. They study cancer to better understand its causes, progression, and consequences in order to therefore eradicate it from the earth more effectively. A distaste for something… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/tactical-military-reasoning-v.1.0.Davis_Square_v1.0
DBbun Davis Square Synthetic Dataset (1800–2200)
A fully synthetic, privacy-free, and educational dataset simulating the evolution of the Davis Square area in Somerville, Massachusetts from the 1800s through the 2200s.
This dataset enables learners, researchers, and developers to explore data science, analytics, and machine learning safely — no real people, addresses, or businesses are represented.
Dataset Summary
Table
Description
geo_streets.csv
Real… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/Davis_Square_v1.0.colossal-oscar-1.0
Dataset Card for Colossal OSCAR 1
IMPORTANT NOTE: THIS DATASET CARD IS STILL BEING WRITTEN, PLEASE BE PATIENT WHILE WE COMPLETE ALL THE INFORMATION ABOUT THE CORPUS
Dataset Summary
The OSCAR project (Open Super-large Crawled Aggregated coRpus) is an Open Source project aiming to provide web-based multilingual resources and datasets for Machine Learning (ML) and Artificial Intelligence (AI) applications. The project focuses specifically in providing large… See the full description on the dataset page: https://huggingface.co/datasets/oscar-corpus/colossal-oscar-1.0.cyberusecase-v1.0
Cybersecurity SOC Fine-Tuning Dataset — 17.5k Real CVEs (2018–2026) + SOC Knowledge
A large supervised fine-tuning (SFT) dataset for teaching an LLM expert-level
cybersecurity reasoning across vulnerability management, SOC alert triage, detection
engineering, threat intelligence & hunting, incident response, and cloud/DevSecOps.
It combines 17,590 real CVEs (2018–2026) pulled from the NIST NVD data feeds with a
hand-curated set of 65 landmark CVEs (rich, multi-angle coverage)… See the full description on the dataset page: https://huggingface.co/datasets/ronaldocloud/cyberusecase-v1.0.stem-reasoning-v1.0.0-ccbysa-001
YouAI Data — stem-reasoning-v1.0.0-ccbysa-001
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 394 step-by-step reasoning chains and 569 instruction/response pairs across 332 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 364 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-v1.0.0-ccbysa-001.ROMB-1.0
♦ ROMB
Русское описание и инструкция
ROMB (Russian Olympiad Math Benchmark) evaluates models on Russian-language
school olympiad mathematics. The test set contains 2552 text-only tasks:
1716 arithmetic/other tasks, 644 logic tasks, and 192 geometry tasks. Tasks have
typed answers, answer-format notes, and per-task checking rules.
The evaluator also supports configurable v3 runs: native thinking, optional
JSON Schema constrained decoding, plain or \boxed{…} answers, and… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/ROMB-1.0.ManipuriGPT-Corpus-v1.0
ManipuriGPT Corpus v1.0
ManipuriGPT Corpus v1.0 is a research-grade, multi-script, deduplicated, and quality-scored corpus specifically engineered for pretraining Manipuri (Meiteilon) language foundation models.
Quick Summary
Total Sequences: 147,956
Total Tokens (ManipuriGPT-Tokenizer-v1.0): 4,347,075
Total Characters: 16,019,401
Pipeline Version: 5.6
Release Version: v1.0.0
Build Timestamp: 2026-07-25T09:17:21.960438Z
Primary Writing Systems… See the full description on the dataset page: https://huggingface.co/datasets/nanskong/ManipuriGPT-Corpus-v1.0.JudgeLM-data-collection-v1.0
Dataset Card for JudgeLM-data-collection
Dataset Summary
This dataset is created for easily use and evaluate JudgeLM. We include LLMs-generated answers and a great multi-modal benchmark, MM-Vet in this repo. The folder structure is shown as bellow:
Folder structure
data
├── JudgeLM/
│ ├── answers/
│ │ ├── alpaca_judgelm_val.jsonl
| | ├── ...
│ ├── judgelm_preprocess.py
│ ├── judgelm_val_5k.jsonl
│ ├── judgelm_val_5k_gpt4.jsonl
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/JudgeLM-data-collection-v1.0.MBG-1.0-data
MBG 1.0 — Dataset
deepRcurs Labs / @deeprcurs · author: Mzed Imamkh / @mzedimamkh
English-only corpora for the MBG 1.0 (Model Bahasa Garuda) project. This is the
external dataset archive; the training/evaluation code lives in the workspace
snapshot (the "controller"), and the model weights live in the model repo
deeprcurs/MBG-1.0.
Contents
File
Rows
Source
Purpose
MBG-1.0.parquet
326,080
generated corpora
main train/val corpus (domain-tagged)… See the full description on the dataset page: https://huggingface.co/datasets/deeprcurs/MBG-1.0-data.qa-expert-multi-hop-qa-V1.0
Dataset Card for QA-Expert-multi-hop-qa-V1.0
This dataset aims to provide multi-domain training data for the task: Question Answering, with a focus on Multi-hop Question Answering.
In total, this dataset contains 25.5k for training and 3.19k for evaluation.
You can take a look at the model we trained on this data: https://huggingface.co/khaimaitien/qa-expert-7B-V1.0
The dataset is mostly generated using the OpenAPI model (gpt-3.5-turbo-instruct). Please read more information about… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/qa-expert-multi-hop-qa-V1.0.JL-ActionBoundary-1K-v1.0.0
JL-ActionBoundary-1K v1.0.0
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code:
ACT: the task is sufficiently specified for bounded repository work;
INSPECT: missing information can be recovered from the repository;
ASK: a material product decision belongs to the user;
DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.OpenOcypus-1.0
🪲 OpenOcypus-1.0
📖 Description
OpenOcypus-1.0 is the first release in the OpenOcypus series — a collection of high‑quality SFT datasets designed to evolve over time. Future versions will introduce additional sources, refined filtering, and expanded task coverage.
Total size: 1,157,428 examples.
The dataset is designed to create a versatile assistant capable of:
🗣️ Engaging in natural conversations
🧮 Solving math problems with step‑by‑step explanations
💻… See the full description on the dataset page: https://huggingface.co/datasets/zait-ai/OpenOcypus-1.0.Tucan-BG-v1.0
Tucan-BG Dataset v1.0
Bilingual Function Calling Training Dataset for Bulgarian Language Models 🇧🇬
📄 Supporting the Tucan model series
Paper: https://arxiv.org/abs/2506.23394
Overview 🚀
Tucan-BG-v1.0 is a bilingual (Bulgarian/English) dataset containing 10,035 conversations specifically designed for training language models in function calling and tool use capabilities. This dataset enables the development of AI agents capable of determining when to use… See the full description on the dataset page: https://huggingface.co/datasets/llm-bg/Tucan-BG-v1.0.
