datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ClimbLab-JaJapanese / 日本語版
ClimbLab-Ja
ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.ClimbLabClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters.
Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.nanochat-climbmix-arithmetic-base10
nanochat ClimbMix + Base-10 Arithmetic
This dataset contains the first 170 shuffled ClimbMix training shards
used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is
mixed into shards 00000..00149; the final
20 train shards are unchanged web-only padding.
The original validation shard (shard_06542.parquet) is also
copied unchanged.
Arithmetic corpus
Family
Examples
a + b = c (all ordered pairs 0..2000, two exposures)
8,008,002
a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.ClimbMixClimbMix is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbMix is a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. It was introduced in this paper.
We proposed a new algorithm to filter and mix the dataset. First, we grouped the data into 1,000 groups based on topic information. Then we applied two classifiers: one to detect advertisements and another to… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbMix.nanochat-climbmix-arithmetic-base7
nanochat ClimbMix + Arithmetic: base-7 numeral world
This is a deterministic base-7 rendering of
Yujivus/nanochat-climbmix-arithmetic-base10. It preserves
the exact shard names, row order, document order, arithmetic-document placement,
and non-numeric text of the source dataset.
Transformation rule
Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10
integer and rendered in base 7. Leading zeros are preserved as a prefix; signs,
punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base7.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.nanochat-climbmix-170
nanochat ClimbMix: first 170 train shards
Convenience mirror of the exact initial ClimbMix slice downloaded by
python -m nanochat.dataset -n 170.
Contents
Training: shard_00000.parquet through shard_00169.parquet
Validation: shard_06542.parquet
manifest.json: pinned source revision, file list, and byte sizes
The Parquet shards are copied without modifying their rows or text.
Attribution and provenance
nanochat:… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-170.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.climbmix-40b-az
ClimbMix 40B — Azerbaijani
A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate.
Dataset Summary
This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language.
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.nanochat-climbmix-arithmetic-base6
nanochat ClimbMix + Arithmetic: base-6 numeral world
This is a deterministic base-6 rendering of
Yujivus/nanochat-climbmix-arithmetic-base10. It preserves
the exact shard names, row order, document order, arithmetic-document placement,
and non-numeric text of the source dataset.
Transformation rule
Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10
integer and rendered in base 6. Leading zeros are preserved as a prefix; signs,
punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base6.Text-to-sql-v1medical-clinical-reasoning-sft-100k
Medical Clinical Reasoning SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education.
Dataset Description
This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.protocolos-clinicos-br
Protocolos Clínicos BR
Paper | Code | Blog post
Brazilian Ministry of Health official clinical guidelines (PCDTs and related) plus the synthetic training corpus derived from them, used to adapt LLMs to Brazilian clinical knowledge. This dataset was introduced in the paper "Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines".
Configurations
default — Original guidelines (raw text)
The 178 official Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/hugo/protocolos-clinicos-br.cli-1m
CLI-1M: Industry-Diverse NL→Shell Training Corpus
975,933 natural-language → shell-command pairs · 18 industries · 6 shells · 13 languages · Apache-2.0
from datasets import load_dataset
ds = load_dataset("carosh/cli-1m", revision="v1.0", split="train")
# 843,461 rows — SFT-ready, license-filtered, quality-gated
The most industry-diverse public dataset for NL→shell-command generation.
108× larger than NL2Bash (the previous public benchmark), and the first
multilingual CLI corpus.… See the full description on the dataset page: https://huggingface.co/datasets/carosh/cli-1m.clinicalbr
ClinicalBr
ClinicalBr is the first bilingual (Portuguese–English) clinical-decision benchmark
built from 2,892 real Brazilian case reports drawn from 36 open-access
medical journals spanning 18 specialties. Every case is provided as a parallel
PT/EN pair and supports four evaluation tasks.
Please refer to the paper for full details on the tasks, methodology, and limitations.
Tasks & metrics
Task
Config
n / lang
Metric
Diagnosis retrieval
diagnosis
2,135… See the full description on the dataset page: https://huggingface.co/datasets/Giordanopsouza/clinicalbr.cli-commands-explained
Overview
This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.ChainSWE
ChainSWE
ChainSWE is a benchmark of sequential, dependent bug fixes for evaluating coding agents on continuous software maintenance. It contains 100 time-ordered chains (304 bug-fix tasks) mined from six SWE-bench-family datasets across 54 Python repositories; each row is one chain over a single repository sharing one base commit and pre-built Docker image, and its bug_fixes field lists the tasks in chronological order, where each task is a self-contained SWE-bench-style… See the full description on the dataset page: https://huggingface.co/datasets/C-lister/ChainSWE.qwen3.5-27b-cli-reasoning-3632x
Qwen3.5-27B CLI Reasoning 3632x
A synthetic reasoning dataset for CLI/terminal command assistance, distilled from Qwen3.5-27B with thinking mode enabled.
Each sample contains a realistic user scenario describing a terminal task, paired with the model's reasoning chain (<think>) and a structured JSON answer (command + description).
Dataset Summary
Source model
Qwen3.5-27B (DashScope API)
Samples
3,632
Thinking mode
Enabled (budget: 4096 tokens)… See the full description on the dataset page: https://huggingface.co/datasets/LocoreMind/qwen3.5-27b-cli-reasoning-3632x.cli-1m
CLI-1M: Industry-Diverse NL→Shell Training Corpus
975,933 natural-language → shell-command pairs · 18 industries · 6 shells · 13 languages · Apache-2.0
from datasets import load_dataset
ds = load_dataset("carosh/cli-1m", revision="v1.0", split="train")
# 843,461 rows — SFT-ready, license-filtered, quality-gated
The most industry-diverse public dataset for NL→shell-command generation.
108× larger than NL2Bash (the previous public benchmark), and the first
multilingual CLI… See the full description on the dataset page: https://huggingface.co/datasets/gavi56/cli-1m.clinical-case-icd10-diagnosis
Clinical History -> ICD-10 (Acute / Chronic) — CC BY-enriched
1798 de-identified clinical histories drawn from open-access case reports in PubMed Central, each paired with a single principal-diagnosis label: an ICD-10-CM code, its official descriptor, and an ACUTE/CHRONIC acuity status.
input: a de-identified clinical history (presentation only; the diagnosis is removed and no PHI is present).
output: {icd10_code, name, status} where status is ACUTE or CHRONIC (how the… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/clinical-case-icd10-diagnosis.devin-cli-reasoning-distillation
Devin CLI Reasoning Distillation Dataset
A distillation dataset built from Devin CLI session traces, containing the model's internal
reasoning traces (chain-of-thought / thinking), user prompts, assistant answers, and tool calls.
The dataset is formatted to be directly compatible with SFT training pipelines that expect
OpenAI-style message lists with a reasoning_content field.
Dataset Summary
Total rows
2,632 (2,507 train / 125 validation)
Rows with… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/devin-cli-reasoning-distillation.clinicaltrial-protocol-corpus
Clinical Trial Protocol Corpus
Full-text clinical trial protocol documents from ClinicalTrials.gov with section segmentation aligned to SPIRIT/ICH-GCP categories.
What this is
49,002 protocol PDFs downloaded from the ClinicalTrials.gov CDN, extracted to text via PyMuPDF, and segmented into structured sections. Each record contains the full protocol text plus a list of detected sections with headings, hierarchy levels, and section type labels from a 15-type… See the full description on the dataset page: https://huggingface.co/datasets/JulesCan/clinicaltrial-protocol-corpus.task685_mmmlu_answer_generation_clinical_knowledge
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task685_mmmlu_answer_generation_clinical_knowledge
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task685_mmmlu_answer_generation_clinical_knowledge.augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/johnny8808/augmented-clinical-notes.clinical-summarizer-sft
Clinical Note Summarizer (SOAP Format)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Long clinical notes → structured SOAP summaries
Why download this
Automate clinical documentation. Reduce physician burnout by summarizing visit notes into Subjective / Objective / Assessment / Plan format.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/clinical-summarizer-sft.psiddx-clinical-ddx
PsiDDx — Clinical Differential-Diagnosis Dataset (v5.2)
A decontaminated, calibrated differential-diagnosis training corpus for the Adaption Labs
AutoScientist Challenge (Healthcare). Each record pairs a patient presentation with a
ranked top-5 differential — common explanations first, each with an ICD-10 code, a calibrated
confidence, and the discriminating feature; rare diagnoses are kept as low-confidence
must-not-miss entries rather than over-called.
5,482 rows. Used to… See the full description on the dataset page: https://huggingface.co/datasets/shariqazeem/psiddx-clinical-ddx.clinical-trials-embeddings
Clinical Trials Embeddings Dataset
Overview
This dataset contains information extracted from clinical trial records collected from ClinicalTrials.gov (Date Accessed: 05/02/2025) along with briefSummary columns embeddings generated using minishlab/potion-base-8M. It focuses on key descriptive fields that provide insight into trial objectives, eligibility criteria, and study design. The dataset is designed for researchers, healthcare professionals, and AI/ML practitioners… See the full description on the dataset page: https://huggingface.co/datasets/cyrilzakka/clinical-trials-embeddings.synthetic_clinical_conversations
Synthetic Clinical Conversations
Fully synthetic English clinical conversations (care-coordination calls, telehealth
visits, post-discharge check-ins) paired with structured encounter records, built for
training and evaluating transcript→JSON extraction models.
Generated structure-first with Tonic Fabricate: the
structured facts are authored as relational data with controlled vocabularies, the
conversation is rendered from those facts, and the extraction target is a… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/synthetic_clinical_conversations.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Vinay393/augmented-clinical-notes.
