datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stratified-kmeans-diverse-pretraining-100K-1M
Stratified K-Means Diverse Pre-Training Dataset (100K-1M)
A carefully balanced subset combining FineWeb-Edu and Proof-Pile-2, featuring embedding-based k-means sampling to ensure diverse representation across educational and mathematical/scientific content at multiple scales.
👥 Follow the Authors
Aman Priyanshu
Supriti Vijay
Overview
This dataset provides stratified subsets at 50k, 100k, 250k, 500k, and 1M scales, combining high-quality… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/stratified-kmeans-diverse-pretraining-100K-1M.leaves-of-grass
leaves of grass
The following is a dataset for training a model to generate text in the style of Walt Whitman's "Leaves of Grass".
The idea with this dataset is to provide a single line (input) and then provide the next lines (1 to 10 lines) of the poem as output.
A model can then be trained to generate lines of poems given a single line of input.
There is a generate_data.py script that can be used to generate the dataset.
It keeps some formatting. New lines may be indented by a… See the full description on the dataset page: https://huggingface.co/datasets/diversen/leaves-of-grass.CodeGen-Diverse-5K
CodeGen-Diverse-5K: Broad Coverage for Competitive Programming
Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset)
Dataset Description
CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions.
Key Statistics
Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.stratified-kmeans-diverse-reasoning-100K-1M
Stratified K-Means Diverse Reasoning Dataset (100K-1M)
A carefully balanced subset of NVIDIA's Llama-Nemotron Post-Training Dataset, featuring square-root rebalanced sampling across math, code, science, instruction-following, chat, and safety tasks at multiple scales.
👥 Follow the Authors
Aman Priyanshu
Supriti Vijay
Overview
This dataset provides stratified subsets at 50k, 100k, 250k, 500k, and 1M scales from the Llama-Nemotron… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/stratified-kmeans-diverse-reasoning-100K-1M.reasoning-sft-minimax-stratified-kmeans-diverse-reasoning-842K-only
MiniMax-M2.5 Reasoning SFT (Stratified K-Means Diverse Reasoning 1M)
Reasoning SFT dataset generated by MiniMaxAI/MiniMax-M2.5 on prompts from the Stratified K-Means Diverse Reasoning 100K-1M dataset.
Format
Each row has three columns:
input — list of dicts [{"role": "...", "content": "..."}, ...] (conversation turns)
response — model-generated response with <think> reasoning block
source — task category (math, code, science, chat, safety)
Generation
Model:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-minimax-stratified-kmeans-diverse-reasoning-842K-only.diverse-svg-prompts
Diverse SVG Prompts
Diverse SVG Prompts is a public collection of 20,000 high-quality,
generated and filtered English briefs for SVG and vector-graphics generation.
It contains 18,000 general illustration prompts and 2,000 lettering prompts.
Schema
The dataset intentionally has only two columns:
prompt: the complete visual brief.
type_tags: a list of category, author-model, and processing tags.
Example:
{
"prompt": "A moonlit mechanical heron..."… See the full description on the dataset page: https://huggingface.co/datasets/Nbardy/diverse-svg-prompts.sinhala-corpus-c-diverse-1m
Diversity-Optimized Sinhala Corpus
A diversity-optimized subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model C) as part of a diversity-driven Sinhala language model adaptation study.
Corpus variants in this series:
Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline)
Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline)… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-c-diverse-1m.stratified-kmeans-diverse-instruction-following-100K-1M
Stratified K-Means Diverse Instruction-Following Dataset (100K-1M)
A carefully balanced subset combining Tulu-3 SFT Mixture and Orca AgentInstruct, featuring embedding-based k-means sampling across diverse instruction-following tasks at multiple scales.
👥 Follow the Authors
Aman Priyanshu
Supriti Vijay
Overview
This dataset provides stratified subsets at 50k, 100k, 250k, 500k, and 1M scales, combining high-quality instruction-following data from… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/stratified-kmeans-diverse-instruction-following-100K-1M.dualmsm-cheese-mixes-diverse
dualmsm-cheese-mixes-diverse
Two finetune-ready cheese-preference mixtures for the dual-MSM cheese dissociation experiments, freshly assembled from the diverse cheese-AFT datasets (the original small sets plus the expanded sets). Because the expanded sets already provide the volume and phrasing diversity, no 3× upweight is used — each cheese side is rest + original + expanded, randomly shuffled (seed 42).
file
rows
teaches
rest_amercheese_diverse.jsonl
29,899
like… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-cheese-mixes-diverse.Diverse-Knowledge
Everything Data
This data is synthetically generated by a ton of open and closed source models. This is basically a parsed version of yearly log form a small dialouge based testing to anylyze model's response on it then perform human evals on it.
The data contains information about everything from every domain, most of the pairs included in this data are preferred by humans as the model's response.
It can be used for topic modeling, or human preference evals etc.
Rest anyone can do… See the full description on the dataset page: https://huggingface.co/datasets/kunu5402/Diverse-Knowledge.BERnaT-Diverse
BERnaT: Basque Encoders for Representing Natural Textual Diversity
Submitted to LREC 2026
Abstract
Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally
exclude non-standard linguistic varieties, reduce model robustness and reinforce representational biases. In this
paper, we argue that language models should aim to capture the full spectrum of language variation (dialectal,
historical, informal, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BERnaT-Diverse.hollow-knight-diverse-dataset
Hollow knight diverse dataset
A dataset containing every tiny detail about the game 'hollow knight'.
date of the wiki data : 5/31/2026
columns : "Question", "Answer"
composition : comprised of Q&A pairs for every page of the hollow knight wiki, all the 557 pages.
reference :all_pages
diverse-2.5m
Diverse Source Text Dataset (2.5M)
A curated, deduplicated, multi-domain English text dataset blending 7 sources across STEM, legal, scientific, encyclopedic, Q&A, and general knowledge domains. Designed as high-quality, diverse source material for downstream NLP tasks such as synthetic data generation, fine-tuning, and text analysis.
Dataset Summary
Total samples
2,500,000
Estimated tokens
~2.8B (GPT-2) / ~2.4B (modern tokenizers)
Language
English… See the full description on the dataset page: https://huggingface.co/datasets/blythet/diverse-2.5m.DeepScaleR-Qwen3-1.7B-2k-diverse-agreed-coded
DeepScaleR-Qwen3-1.7B-2k diverse-agreed, strategy-coded
1635 competition-math problems (the claude_agrees_gold == True subset of a
2k diverse-classified DeepScaleR pool). Each row carries Claude's worked
claude_solution plus three leak-free re-expressions of the strategy it
deploys, drawn from a shared 116-code strategy codebook.
Columns
idx — row index into agentica-org/DeepScaleR-Preview-Dataset (resume/join key).
problem, answer — the problem and gold answer.… See the full description on the dataset page: https://huggingface.co/datasets/zjhhhh/DeepScaleR-Qwen3-1.7B-2k-diverse-agreed-coded.diverse_sokoban
Debunk the Myth of SFT Generalization Dataset
This dataset is associated with the paper Debunk the Myth of SFT Generalization, which re-evaluates the generalization capabilities of supervised fine-tuning (SFT) compared to reinforcement learning (RL) on decision-making benchmarks. The research demonstrates that with proper data curation, such as prompt diversity and Chain-of-Thought (CoT) supervision, SFT can achieve strong generalization, matching or even surpassing RL baselines.… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse_sokoban.diverse-answer-only-gp-l-only-10k
General Points Dataset from Debunk the Myth of SFT Generalization
This dataset is part of the research presented in the paper Debunk the Myth of SFT Generalization. It contains data for the General Points decision-making benchmark, which is used to evaluate the generalization capabilities of Supervised Fine-Tuning (SFT) models against Reinforcement Learning (RL) baselines. The paper explores the impact of prompt diversity and Chain-of-Thought (CoT) supervision on SFT's ability to… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse-answer-only-gp-l-only-10k.diverse-answer-only-sokoban
Dataset from "Debunk the Myth of SFT Generalization"
This dataset is associated with the research presented in the paper Debunk the Myth of SFT Generalization.
The paper challenges the conventional wisdom that supervised fine-tuning (SFT) primarily memorizes training data and struggles with generalization, contrasting it with reinforcement learning (RL)'s perceived robustness. Through systematic evaluation on decision-making benchmarks such as Sokoban and General Points, the… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse-answer-only-sokoban.diverse-cot-gp-l-only-10k
General Points Dataset from "Debunk the Myth of SFT Generalization"
This dataset is part of the research presented in the paper "Debunk the Myth of SFT Generalization". The paper challenges the narrative that supervised fine-tuning (SFT) is inherently inferior to reinforcement learning (RL) by demonstrating SFT's strong generalization capabilities on decision-making benchmarks.
This particular dataset focuses on the "General Points" task, which involves arithmetic with five-card… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse-cot-gp-l-only-10k.diverse-websearch-3.5k
Diverse WebSearch 3.5k
Diverse WebSearch 3.5K is a small web research dataset containing 3,500 diverse web pages collected from search results. Each row contains a source URL, extracted markdown content, a concise summary, and image URLs found on the page.
Dataset Details
This dataset is intended for learning and experimentation with:
webpage summarization
retrieval-augmented generation
search result understanding
document cleaning
synthetic QA generation
dataset… See the full description on the dataset page: https://huggingface.co/datasets/soumikmahato/diverse-websearch-3.5k.diverse-cot-sokoban
Debunk the Myth of SFT Generalization Dataset
This dataset is part of the research presented in the paper Debunk the Myth of SFT Generalization.
A prevailing view holds that supervised fine-tuning (SFT) memorizes training data and fails to generalize, whereas reinforcement learning (RL) attains broader robustness. This paper challenges this claim through a systematic evaluation on decision-making benchmarks, Sokoban and General Points. It shows that much of SFT's perceived failure… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse-cot-sokoban.diverse-bash-dataset-extended
Bash diverse dataset extended
This dataset is a continious work from datasetter458/bash-diverse-dataset
Commands covered: bash, apt, crontab, date, diff, docker, env, export, fdisk, file, fsck, gcc, git, history, join, kubectl, lsof, man, makefs, mount, nice, node, nohup, npm,
passwd, paste, patch, pkill, python, renice, split, ss, time, umount, uname, uniq, vi, whereis, which
Note : For each command, a big amount of Q&As, and each set of them covers pretty much all the flags for… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/diverse-bash-dataset-extended.loracle-ia-diverse-qa-subagent-10q
Loracle IA Diverse QA Subagent 10Q
This dataset is a derived, expanded version of ceselder/loracle-ia-diverse-qa.
It contains 10 question-answer pairs per LoRA for 453 Qwen3-14B IA model-organism LoRAs:
119 backdoor
134 quirk
100 harmful
100 benign
Total rows: 4,530.
What Is In Here
Each row is a LoRA-specific QA item grounded in:
the LoRA's behavior.txt
two selected support prompts from its train.jsonl
a same-family distractor LoRA
a paired mirror LoRA when… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracle-ia-diverse-qa-subagent-10q.diverse-bash-dataset
Diverse bash terminal dataset
A dataset containing Question&Code pairs for most of the standard bash commands.
columns : "Question", "Code answer".
Commands covered : echo, cat, cd, rm, mkdir, top, free, du, df, ps, head, tail, grep, cp, cut, sort, touch, ls, groupadd, ifconfig, ip, ln, ping, scp, ssh, sudo, systemctl, tar, useradd, userdel, usermod, wc
specs : high quality dataset with precise coding examples for each command, covering (almost) all the flags of every command… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/diverse-bash-dataset.
