datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-ClimbLab
ClimbLab Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbLab.ClimbLab-JaJapanese / 日本語版
ClimbLab-Ja
ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.ClimbLabClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters.
Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.Nemotron-ClimbMix
ClimbMix Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.nanochat-climbmix-arithmetic-base10
nanochat ClimbMix + Base-10 Arithmetic
This dataset contains the first 170 shuffled ClimbMix training shards
used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is
mixed into shards 00000..00149; the final
20 train shards are unchanged web-only padding.
The original validation shard (shard_06542.parquet) is also
copied unchanged.
Arithmetic corpus
Family
Examples
a + b = c (all ordered pairs 0..2000, two exposures)
8,008,002
a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.ClimbMixClimbMix is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbMix is a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. It was introduced in this paper.
We proposed a new algorithm to filter and mix the dataset. First, we grouped the data into 1,000 groups based on topic information. Then we applied two classifiers: one to detect advertisements and another to… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbMix.nanochat-climbmix-arithmetic-base7
nanochat ClimbMix + Arithmetic: base-7 numeral world
This is a deterministic base-7 rendering of
Yujivus/nanochat-climbmix-arithmetic-base10. It preserves
the exact shard names, row order, document order, arithmetic-document placement,
and non-numeric text of the source dataset.
Transformation rule
Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10
integer and rendered in base 7. Leading zeros are preserved as a prefix; signs,
punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base7.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.nanochat-climbmix-170
nanochat ClimbMix: first 170 train shards
Convenience mirror of the exact initial ClimbMix slice downloaded by
python -m nanochat.dataset -n 170.
Contents
Training: shard_00000.parquet through shard_00169.parquet
Validation: shard_06542.parquet
manifest.json: pinned source revision, file list, and byte sizes
The Parquet shards are copied without modifying their rows or text.
Attribution and provenance
nanochat:… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-170.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.ClimbMix-6BT
ClimbMix-6BT
This is the tokenized nvidia/Nemotron-ClimbMix (10M subset) using SmolLM2-135M tokenzier. Data is divided into shards (.npy files) for easier to load with PyTorch IterableDataset.
Each .npy file can be loaded with numpy.load('file_name.npy').
Split
# Documents
# Shards
# Tokens
train
9,900,000
65
6,463,974,020 (6.5B)
val
100,000
1
64,859,672 (65M)
Total
10,000,000
66
6,528,833,692 (6.5B)
Example of usage
uvx hf download… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/ClimbMix-6BT.climbmix-tokenized-20480-diloco
ClimbMix, retokenized and shuffled for three-worker DiLoCo
This is a document-preserving, three-way split of NVIDIA's
Nemotron-ClimbMix,
retokenized with a 20,480-entry byte-level BPE tokenizer. Each document ends in
<|endoftext|>. The Arrow IPC streams use transparent Zstandard buffer
compression. A deterministic whole-shard holdout is shared by every worker for
validation and is excluded from training.
Training part
Documents
Tokens
Files
Compressed size
000
15,709… See the full description on the dataset page: https://huggingface.co/datasets/Sambarboi/climbmix-tokenized-20480-diloco.climbmix-40b-az
ClimbMix 40B — Azerbaijani
A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate.
Dataset Summary
This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language.
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.nanochat-climbmix-arithmetic-base6
nanochat ClimbMix + Arithmetic: base-6 numeral world
This is a deterministic base-6 rendering of
Yujivus/nanochat-climbmix-arithmetic-base10. It preserves
the exact shard names, row order, document order, arithmetic-document placement,
and non-numeric text of the source dataset.
Transformation rule
Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10
integer and rendered in base 6. Leading zeros are preserved as a prefix; signs,
punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base6.Text-to-sql-v1protocolos-clinicos-br
Protocolos Clínicos BR
Paper | Code | Blog post
Brazilian Ministry of Health official clinical guidelines (PCDTs and related) plus the synthetic training corpus derived from them, used to adapt LLMs to Brazilian clinical knowledge. This dataset was introduced in the paper "Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines".
Configurations
default — Original guidelines (raw text)
The 178 official Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/hugo/protocolos-clinicos-br.medical-clinical-reasoning-sft-100k
Medical Clinical Reasoning SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education.
Dataset Description
This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.cli-1m
CLI-1M: Industry-Diverse NL→Shell Training Corpus
975,933 natural-language → shell-command pairs · 18 industries · 6 shells · 13 languages · Apache-2.0
from datasets import load_dataset
ds = load_dataset("carosh/cli-1m", revision="v1.0", split="train")
# 843,461 rows — SFT-ready, license-filtered, quality-gated
The most industry-diverse public dataset for NL→shell-command generation.
108× larger than NL2Bash (the previous public benchmark), and the first
multilingual CLI corpus.… See the full description on the dataset page: https://huggingface.co/datasets/carosh/cli-1m.climbmix-seed42-10b-replay
ClimbMix seed-42 training replay, approximately 10B tokens
manifest.json is the authoritative export status: only state: ready means
construction is complete. It has not been uploaded to Hugging Face.
This is a frozen training input stream, not a new raw-text mixture. It uses
the existing OptimalScale/ClimbMix revision and the exact tokenizer pinned in
the manifest. Documents are shuffled with seed 42 and split into the original
eight contiguous virtual-rank streams before… See the full description on the dataset page: https://huggingface.co/datasets/ArchloverLRZ/climbmix-seed42-10b-replay.code_clippyThis dataset was generated by selecting GitHub repositories from a large collection of repositories. These repositories were collected from https://seart-ghs.si.usi.ch/ and Github portion of [The Pile](https://github.com/EleutherAI/github-downloader) (performed on July 7th, 2021). The goal of this dataset is to provide a training set for pretraining large language models on code data for helping software engineering researchers better understand their impacts on software related tasks such as autocompletion of code. The dataset is split into train, validation, and test splits. There is a version containing duplicates (209GBs compressed) and ones where exact duplicates (132GBs compressed) are removed. Contains mostly JavaScript and Python code, but other programming languages are included as well to various degrees.cli-commands-explained
Overview
This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.clinicalbr
ClinicalBr
ClinicalBr is the first bilingual (Portuguese–English) clinical-decision benchmark
built from 2,892 real Brazilian case reports drawn from 36 open-access
medical journals spanning 18 specialties. Every case is provided as a parallel
PT/EN pair and supports four evaluation tasks.
Please refer to the paper for full details on the tasks, methodology, and limitations.
Tasks & metrics
Task
Config
n / lang
Metric
Diagnosis retrieval
diagnosis
2,135… See the full description on the dataset page: https://huggingface.co/datasets/Giordanopsouza/clinicalbr.ChainSWE
ChainSWE
ChainSWE is a benchmark of sequential, dependent bug fixes for evaluating coding agents on continuous software maintenance. It contains 100 time-ordered chains (304 bug-fix tasks) mined from six SWE-bench-family datasets across 54 Python repositories; each row is one chain over a single repository sharing one base commit and pre-built Docker image, and its bug_fixes field lists the tasks in chronological order, where each task is a self-contained SWE-bench-style… See the full description on the dataset page: https://huggingface.co/datasets/C-lister/ChainSWE.qwen3.5-27b-cli-reasoning-3632x
Qwen3.5-27B CLI Reasoning 3632x
A synthetic reasoning dataset for CLI/terminal command assistance, distilled from Qwen3.5-27B with thinking mode enabled.
Each sample contains a realistic user scenario describing a terminal task, paired with the model's reasoning chain (<think>) and a structured JSON answer (command + description).
Dataset Summary
Source model
Qwen3.5-27B (DashScope API)
Samples
3,632
Thinking mode
Enabled (budget: 4096 tokens)… See the full description on the dataset page: https://huggingface.co/datasets/LocoreMind/qwen3.5-27b-cli-reasoning-3632x.cli-1m
CLI-1M: Industry-Diverse NL→Shell Training Corpus
975,933 natural-language → shell-command pairs · 18 industries · 6 shells · 13 languages · Apache-2.0
from datasets import load_dataset
ds = load_dataset("carosh/cli-1m", revision="v1.0", split="train")
# 843,461 rows — SFT-ready, license-filtered, quality-gated
The most industry-diverse public dataset for NL→shell-command generation.
108× larger than NL2Bash (the previous public benchmark), and the first
multilingual CLI… See the full description on the dataset page: https://huggingface.co/datasets/gavi56/cli-1m.clinical-case-icd10-diagnosis
Clinical History -> ICD-10 (Acute / Chronic) — CC BY-enriched
1798 de-identified clinical histories drawn from open-access case reports in PubMed Central, each paired with a single principal-diagnosis label: an ICD-10-CM code, its official descriptor, and an ACUTE/CHRONIC acuity status.
input: a de-identified clinical history (presentation only; the diagnosis is removed and no PHI is present).
output: {icd10_code, name, status} where status is ACUTE or CHRONIC (how the… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/clinical-case-icd10-diagnosis.devin-cli-reasoning-distillation
Devin CLI Reasoning Distillation Dataset
A distillation dataset built from Devin CLI session traces, containing the model's internal
reasoning traces (chain-of-thought / thinking), user prompts, assistant answers, and tool calls.
The dataset is formatted to be directly compatible with SFT training pipelines that expect
OpenAI-style message lists with a reasoning_content field.
Dataset Summary
Total rows
2,632 (2,507 train / 125 validation)
Rows with… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/devin-cli-reasoning-distillation.clinicaltrial-protocol-corpus
Clinical Trial Protocol Corpus
Full-text clinical trial protocol documents from ClinicalTrials.gov with section segmentation aligned to SPIRIT/ICH-GCP categories.
What this is
49,002 protocol PDFs downloaded from the ClinicalTrials.gov CDN, extracted to text via PyMuPDF, and segmented into structured sections. Each record contains the full protocol text plus a list of detected sections with headings, hierarchy levels, and section type labels from a 15-type… See the full description on the dataset page: https://huggingface.co/datasets/JulesCan/clinicaltrial-protocol-corpus.task685_mmmlu_answer_generation_clinical_knowledge
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task685_mmmlu_answer_generation_clinical_knowledge
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task685_mmmlu_answer_generation_clinical_knowledge.augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.
