datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
physical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.megawika-report-generation
Dataset Card for MegaWika for Report Generation
Dataset Summary
MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span
50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a
non-English language, an automated English translation is provided.
This dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/megawika-report-generation.200k_HEAVY_gpt4o-description-gpt4omini-code_generated_problemsHere is the dataset of ~100k synthetic data generated by 162 seeds.
We generate the dataset with the following steps and two approaches:
Generate ~110k descriptions by GPT4o.
Approach 1: Generate ~110k codes follow each description by GPT4o-mini.
Approach 2: Generate ~110k codes follow each description by GPT4o-mini and suggest it to use specific library functions.
Run the ~220k codes and do auto-filtering.
Get the final ~200k legitimate ARC-like tasks with examples.
General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.GeneratingQuestions
HVU_QA
HVU_QA is an open-source Vietnamese Question-Context-Answer (QCA) corpus, accompanied by supporting tools, created to facilitate the development of FAQ-style question generation and question answering systems, particularly for low-resource language settings. The dataset was developed by a research team at Hung Vuong University, Phu Tho, Vietnam, led by Dr. Ha Nguyen, Deputy Head of the Department of Engineering Technology. HVU_QA was constructed using a fully automated… See the full description on the dataset page: https://huggingface.co/datasets/DANGDOCAO/GeneratingQuestions.General-Stories-CollectionGeneral Stories Collection
A great synthetic datasets consists of around 1.3 million stories especially meant for General audience. You can directly use these datasets for training large models.
Total 10 datasets are available for download. You can use any one or all the json files for training purpose.
These datasets are in "prompt" and "text" format. Total token length is also available.
Thanks for your love & support.
week1-general-20b-dolma2-v1
Week-One General 20B Dolma2
This is a deterministic, pretokenized 20-billion-token baseline corpus for
controlled language-model architecture and training experiments. It contains
nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered
extension of the previous view. It also includes a dataset-only
370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B
order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.GEN
Human, AI-Generated, and AI-Edited Text: Stylometric Corpus
📄 Paper: https://arxiv.org/pdf/2608.27855
💻 Code: https://github.com/ZhengyangShan/GEN-stylometric-footprint
A three-class corpus for studying how AI writing differs from human writing,
distinguishing two modes of AI involvement:
AI generation: text written by an LLM from scratch, given a prompt.
AI editing: human text revised by an LLM (grammar, tone, paraphrase, etc.).
Supports detection of AI-generated text… See the full description on the dataset page: https://huggingface.co/datasets/szyszy/GEN.SWE-Skills-BenchDataset Summary
SWE-Skills-Bench is a benchmark dataset for evaluating whether injected skill documents — structured packages of procedural knowledge — measurably improve LLM agent performance on real-world software engineering tasks.
The dataset contains 49 skills spanning 565 task instances across six software engineering domains (Deployment & DevOps, Analytics & Monitoring, API Development, Data Science & ML, Security & Testing, and Developer Tools). Each skill is grounded in an authentic… See the full description on the dataset page: https://huggingface.co/datasets/GeniusHTX/SWE-Skills-Bench.backend-code-generator-dataset
Backend Code Generation Dataset
Dataset Description
This dataset contains examples for training AI models to generate backend application code. It includes descriptions of backend requirements paired with complete, functional code implementations across multiple frameworks and programming languages.
Dataset Summary
The Backend Code Generation Dataset is designed to train models that can generate complete backend applications from natural language descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/Techta/backend-code-generator-dataset.Nemotron-RLHF-GenRM-v1
Dataset Description:
This dataset is designed to train Generative Reward Models (GenRMs). It leverages reinforcement learning at scale to train accurate and robust GenRMs that generalize better than traditional Bradley-Terry models and reduce the risk of reward hacking.
The dataset is composed of:
Preference data focused on diverse domains
A synthetic safety blend
The data follows a "meta-prompt" structure where the model is instructed to act as an expert evaluation judge. For… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1.SFT-General-Japanese-60K
SFT-General-Japanese-60K
Welcome to this dataset! 👋
Need clean, natural Japanese conversations for supervised fine-tuning? You are in the right place. SFT-General-Japanese-60K contains 60,000 carefully filtered instruction–response conversations ready for chat-model training. It combines the practical breadth of open Japanese SFT data with transparent gates for safety, recency, formatting, language consistency, and redundancy—so you can focus on training rather… See the full description on the dataset page: https://huggingface.co/datasets/OysterCoreAI/SFT-General-Japanese-60K.100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4o-mini.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
SFT-General-Spanish-50K
SFT-General-Spanish-50K now available🎇
Una buena conversación no necesita hacer ruido: necesita entender la pregunta, ordenar lo importante y dejar a la otra persona con un siguiente paso claro. SFT-General-Spanish-50K reúne 50.413 conversaciones originales en español diseñadas para entrenar ese tipo de ayuda.
Resumen
Registros: 50.413 (no se redondeó a 50K).
Idioma: español contemporáneo, registro general y neutro.
Formato: JSONL de mensajes estilo chat; tres… See the full description on the dataset page: https://huggingface.co/datasets/OysterCoreAI/SFT-General-Spanish-50K.Aloe-Beta-General-Collection
Aloe-Beta-Medical-Collection
Collection of curated general datasets used to fine-tune Aloe-Beta.
Dataset Details
Dataset Description
We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including:
Coding, math, data analysis, STEM, etc.
Function calling
Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
generation-ship-world
Generation Ship — Multi-AI Collaborative Future History (2025–3000+)
A 1,000-year future history whose canon is written by AI agents. Hard rules, archival fiction,
no omniscient narration. 13 artifacts from 5 LLMs so far (claude-sonnet-5, gpt-5, minimax-m3,
deepseek-v4-pro, gemini-3.7-flash).
Contents
Path
What it is
core/世界规则.md
The world's hard rules: physics (no FTL, no cryosleep, 0.03c fusion-pulse ship, 200 years to Proxima b), history (7 eras… See the full description on the dataset page: https://huggingface.co/datasets/dahongge/generation-ship-world.AD-GEN
AD-GEN: Evidence-Preserving Generation of Validated ATT&CK-Aligned Narratives from Large-Scale Endpoint Telemetry
LLM-ready SOC narratives derived from large-scale Windows endpoint telemetry
Overview
Modern endpoint telemetry datasets contain rich behavioral evidence but are often difficult to use directly for large language model (LLM) reasoning. Raw Sysmon logs are fragmented across individual events, contain sensitive identifiers, suffer from process… See the full description on the dataset page: https://huggingface.co/datasets/namhop88/AD-GEN.test2
T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains
T1-Bench is a high-fidelity benchmark for evaluating task-completion and role-playing agents across 25 domains, including 11 single-domain and 14 multi-domain settings. It provides 76 tools and extensive human annotations, enabling systematic evaluation of agents in realistic, policy-grounded multi-domain interactions with natural user–assistant role-playing.
T1-Bench is a fully automated benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/gentaiscool/test2.NeuroBio-GenZ-1K
NeuroBio GenZ 1K
Around 1000 neuroscience and biology questions, answered like your smartest friend is texting you back, not like a textbook is talking at you.
"Why does doomscrolling give me dopamine?" gets answered in three sentences, casual tone, real neuroscience terms (nucleus accumbens, not "reward center"), zero fluff.
Why did you make this?
Because there's genuinely not that much high quality neuroscience and biology data on Hugging Face that isn't either… See the full description on the dataset page: https://huggingface.co/datasets/luka0x12/NeuroBio-GenZ-1K.glm5.2-general-distill
Teacher-generated instruction/response pairs used to distill small, local student models
(the ADI / Advanced Data Intelligence series) from the frontier teacher glm-5.2.
How it was built
Teacher: glm-5.2 (served via Ollama Cloud as glm-5.2:cloud), queried with
thinking/reasoning disabled so every record is a single clean final answer.
Seed prompts: databricks/databricks-dolly-15k,
filtered to remove items that require an attached context passage — the closed_qa… See the full description on the dataset page: https://huggingface.co/datasets/AdvancedDataIntelligence/glm5.2-general-distill.qwen-generated-svamp-controls-sft
Qwen-Generated SVAMP CoT Controls ? SFT
Qwen-generated controlled reasoning traces for SVAMP in LLaMA-Factory SFT format. Variants include ordinary, all-caps, no-comma, disclaimer, and multilingual examples.
Splits
3,940 training examples and 380 held-out evaluation examples.
Format
The JSON files use the LLaMA-Factory Alpaca-style schema. The included
dataset_info.json registers the exact training and evaluation names. DPO
records are marked with… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-svamp-controls-sft.gg_zh_v1_550BCCI-Data
SkyPile-150B
TeleChat-PTD
WebText-cn
WuDaoCorpus2.0
wangan
yayi2_pretrain_data
整合+minhash去重了一波,最终得到550B中文预训练语料
TSD-KD-Qwen2.5-1.5B-Instruct-Gen
TSD-KD-Qwen2.5-1.5B-Instruct-Gen
This dataset contains student-generated examples used for Token-Selective Dual Knowledge Distillation (TSD-KD), introduced in our ICLR 2026 paper:
"Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation"
Paper: https://arxiv.org/abs/2603.13260
Github: https://github.com/kmswin1/TSD-KD
Dataset Description
This dataset contains student-generated instruction-response examples from… See the full description on the dataset page: https://huggingface.co/datasets/Minsang/TSD-KD-Qwen2.5-1.5B-Instruct-Gen.SWE-QA-Benchmark
SWE-QA Benchmark
A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories.
Dataset Summary
Total Questions: 720
Repositories: 15
Format: JSONL (JSON Lines)
Fields: question, answer
Repository Coverage
Each repository contains 48 questions:
astropy
conan
django
flask
matplotlib
pylint
pytest
reflex
requests
scikit-learn
sphinx
sqlfluff
streamlink
sympy
xarray… See the full description on the dataset page: https://huggingface.co/datasets/GeniusWondering/SWE-QA-Benchmark.dataset-ohada-droit-commercial-general-echantillon
Dataset OHADA — Droit Commercial Général (AUDCG) — Échantillon
Description
Échantillon de 10 entrées extraites d'un dataset de fine-tuning juridique en cours de conception, portant sur l'Acte Uniforme relatif au Droit Commercial Général (AUDCG) — le texte fondamental du statut du commerçant, des actes de commerce, de la preuve et de la prescription en matière commerciale dans l'espace OHADA (Organisation pour l'Harmonisation en Afrique du Droit des Affaires — 17… See the full description on the dataset page: https://huggingface.co/datasets/Bisilivan/dataset-ohada-droit-commercial-general-echantillon.ReasonFlux_SFT_15k
ReasonFlux_SFT_15k
This dataset contains 15,000 Chinese competition-level training examples for the GaoKao benchmark. It's part of the ReasonFlux-Zero project, which aims to improve LLM reasoning capabilities through a hierarchical reinforcement learning algorithm and a library of thought templates. This dataset was used in the Supervised Fine-Tuning (SFT) stage of ReasonFlux-Zero's development.
Arxiv: https://arxiv.org/abs/2502.06772
Github:… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/ReasonFlux_SFT_15k.gujarati-general-purpose-instruction
Gujarati General-Purpose Instruction Dataset (GGJI v1)
Dataset Summary
GGJI v1 (Gujarati General-Purpose Instruction v1) is a large-scale, high-quality supervised fine-tuning (SFT) dataset designed to train instruction-following language models in Gujarati. It contains 23,181 records across 18 behavioral task categories, covering a broad range of NLP tasks including question answering, summarization, translation, reasoning, creative writing, code explanation, and… See the full description on the dataset page: https://huggingface.co/datasets/tkdonda/gujarati-general-purpose-instruction.Skill2-Bench
Skill²-Bench
Skill²-Bench is a benchmark of multi-step tasks that force LLMs to switch between skills, introduced in the paper "Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning".
Long-horizon tasks require models to switch between skills, not just execute a single skill well. Each Skill²-Bench task embeds a sequence of 2–10 steps in a coherent real-world scenario, where consecutive steps draw on different skills (e.g., algorithm design… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/Skill2-Bench.code-generation-sft-100k
Code Generation SFT (100K)
100,000 ShareGPT conversations covering code generation across 8 programming languages, 21 categories, and 22 distinct programming tasks. Each example includes a detailed natural language request and a complete, working implementation with explanations of key design decisions.
Motivation
Coding assistants are the highest-adoption LLM application category, but most open training datasets focus on isolated functions without context. This… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-generation-sft-100k.
