CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-ClimbLab ClimbLab Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbLab.text-generation1B<n<10B38 likes12k downloads1y agoHugging Face02KantaHayashiAI /ClimbLab-JaJapanese / 日本語版 ClimbLab-Ja ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.tabulartext-generation100M<n<1B2 likes10k downloads4mo agoHugging Face03OptimalScale /ClimbLabClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters. Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.texttext-generation1B<n<10B16 likes9.8k downloads1y agoHugging Face04nvidia /Nemotron-ClimbMix ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.tabulartext-generation100M<n<1B128 likes6.6k downloads11mo agoHugging Face05Yujivus /nanochat-climbmix-arithmetic-base10 nanochat ClimbMix + Base-10 Arithmetic This dataset contains the first 170 shuffled ClimbMix training shards used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is mixed into shards 00000..00149; the final 20 train shards are unchanged web-only padding. The original validation shard (shard_06542.parquet) is also copied unchanged. Arithmetic corpus Family Examples a + b = c (all ordered pairs 0..2000, two exposures) 8,008,002 a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.texttext-generation10M<n<100M0 likes3.8k downloads1mo agoHugging Face06OptimalScale /ClimbMixClimbMix is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbMix is a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. It was introduced in this paper. We proposed a new algorithm to filter and mix the dataset. First, we grouped the data into 1,000 groups based on topic information. Then we applied two classifiers: one to detect advertisements and another to… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbMix.tabulartext-generation100M<n<1B36 likes3.3k downloads1y agoHugging Face07Yujivus /nanochat-climbmix-arithmetic-base7 nanochat ClimbMix + Arithmetic: base-7 numeral world This is a deterministic base-7 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 7. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base7.texttext-generation10M<n<100M0 likes1.5k downloads1mo agoHugging Face08AGBonnet /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.texttext-generation10K<n<100K75 likes1.3k downloads3y agoHugging Face09Yujivus /nanochat-climbmix-170 nanochat ClimbMix: first 170 train shards Convenience mirror of the exact initial ClimbMix slice downloaded by python -m nanochat.dataset -n 170. Contents Training: shard_00000.parquet through shard_00169.parquet Validation: shard_06542.parquet manifest.json: pinned source revision, file list, and byte sizes The Parquet shards are copied without modifying their rows or text. Attribution and provenance nanochat:… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-170.texttext-generation10M<n<100M0 likes928 downloads1mo agoHugging Face10starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes890 downloads2y agoHugging Face11minhnguyent546 /ClimbMix-6BT ClimbMix-6BT This is the tokenized nvidia/Nemotron-ClimbMix (10M subset) using SmolLM2-135M tokenzier. Data is divided into shards (.npy files) for easier to load with PyTorch IterableDataset. Each .npy file can be loaded with numpy.load('file_name.npy'). Split # Documents # Shards # Tokens train 9,900,000 65 6,463,974,020 (6.5B) val 100,000 1 64,859,672 (65M) Total 10,000,000 66 6,528,833,692 (6.5B) Example of usage uvx hf download… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/ClimbMix-6BT.text-generation0 likes744 downloads6mo agoHugging Face12Sambarboi /climbmix-tokenized-20480-diloco ClimbMix, retokenized and shuffled for three-worker DiLoCo This is a document-preserving, three-way split of NVIDIA's Nemotron-ClimbMix, retokenized with a 20,480-entry byte-level BPE tokenizer. Each document ends in <|endoftext|>. The Arrow IPC streams use transparent Zstandard buffer compression. A deterministic whole-shard holdout is shared by every worker for validation and is excluded from training. Training part Documents Tokens Files Compressed size 000 15,709… See the full description on the dataset page: https://huggingface.co/datasets/Sambarboi/climbmix-tokenized-20480-diloco.text-generation1M<n<10M0 likes631 downloads2mo agoHugging Face13LocalDoc /climbmix-40b-az ClimbMix 40B — Azerbaijani A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate. Dataset Summary This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language. Property Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.texttext-generation10M<n<100M1 likes626 downloads6mo agoHugging Face14Yujivus /nanochat-climbmix-arithmetic-base6 nanochat ClimbMix + Arithmetic: base-6 numeral world This is a deterministic base-6 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 6. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base6.texttext-generation10M<n<100M0 likes579 downloads1mo agoHugging Face15Clinton /Text-to-sql-v1texttext-generation100K<n<1M73 likes355 downloads3y agoHugging Face16hugo /protocolos-clinicos-br Protocolos Clínicos BR Paper | Code | Blog post Brazilian Ministry of Health official clinical guidelines (PCDTs and related) plus the synthetic training corpus derived from them, used to adapt LLMs to Brazilian clinical knowledge. This dataset was introduced in the paper "Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines". Configurations default — Original guidelines (raw text) The 178 official Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/hugo/protocolos-clinicos-br.texttext-generation10K<n<100K0 likes322 downloads2mo agoHugging Face17stindardlogic /medical-clinical-reasoning-sft-100k Medical Clinical Reasoning SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education. Dataset Description This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.texttext-generation100K<n<1M0 likes306 downloads2mo agoHugging Face18carosh /cli-1m CLI-1M: Industry-Diverse NL→Shell Training Corpus 975,933 natural-language → shell-command pairs · 18 industries · 6 shells · 13 languages · Apache-2.0 from datasets import load_dataset ds = load_dataset("carosh/cli-1m", revision="v1.0", split="train") # 843,461 rows — SFT-ready, license-filtered, quality-gated The most industry-diverse public dataset for NL→shell-command generation. 108× larger than NL2Bash (the previous public benchmark), and the first multilingual CLI corpus.… See the full description on the dataset page: https://huggingface.co/datasets/carosh/cli-1m.texttext-generation1M<n<10M1 likes229 downloads4mo agoHugging Face19ArchloverLRZ /climbmix-seed42-10b-replay ClimbMix seed-42 training replay, approximately 10B tokens manifest.json is the authoritative export status: only state: ready means construction is complete. It has not been uploaded to Hugging Face. This is a frozen training input stream, not a new raw-text mixture. It uses the existing OptimalScale/ClimbMix revision and the exact tokenizer pinned in the manifest. Documents are shuffled with seed 42 and split into the original eight contiguous virtual-rank streams before… See the full description on the dataset page: https://huggingface.co/datasets/ArchloverLRZ/climbmix-seed42-10b-replay.text-generation1M<n<10M0 likes202 downloads5d agoHugging Face20CodedotAI /code_clippyThis dataset was generated by selecting GitHub repositories from a large collection of repositories. These repositories were collected from https://seart-ghs.si.usi.ch/ and Github portion of [The Pile](https://github.com/EleutherAI/github-downloader) (performed on July 7th, 2021). The goal of this dataset is to provide a training set for pretraining large language models on code data for helping software engineering researchers better understand their impacts on software related tasks such as autocompletion of code. The dataset is split into train, validation, and test splits. There is a version containing duplicates (209GBs compressed) and ones where exact duplicates (132GBs compressed) are removed. Contains mostly JavaScript and Python code, but other programming languages are included as well to various degrees.text-generation12 likes178 downloads4y agoHugging Face21b-mc2 /cli-commands-explained Overview This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.tabulartext-generation10K<n<100K5 likes170 downloads2y agoHugging Face22Giordanopsouza /clinicalbr ClinicalBr ClinicalBr is the first bilingual (Portuguese–English) clinical-decision benchmark built from 2,892 real Brazilian case reports drawn from 36 open-access medical journals spanning 18 specialties. Every case is provided as a parallel PT/EN pair and supports four evaluation tasks. Please refer to the paper for full details on the tasks, methodology, and limitations. Tasks & metrics Task Config n / lang Metric Diagnosis retrieval diagnosis 2,135… See the full description on the dataset page: https://huggingface.co/datasets/Giordanopsouza/clinicalbr.textquestion-answering10K<n<100K0 likes156 downloads29d agoHugging Face23C-lister /ChainSWE ChainSWE ChainSWE is a benchmark of sequential, dependent bug fixes for evaluating coding agents on continuous software maintenance. It contains 100 time-ordered chains (304 bug-fix tasks) mined from six SWE-bench-family datasets across 54 Python repositories; each row is one chain over a single repository sharing one base commit and pre-built Docker image, and its bug_fixes field lists the tasks in chronological order, where each task is a self-contained SWE-bench-style… See the full description on the dataset page: https://huggingface.co/datasets/C-lister/ChainSWE.texttext-generationn<1K0 likes149 downloads3mo agoHugging Face24LocoreMind /qwen3.5-27b-cli-reasoning-3632x Qwen3.5-27B CLI Reasoning 3632x A synthetic reasoning dataset for CLI/terminal command assistance, distilled from Qwen3.5-27B with thinking mode enabled. Each sample contains a realistic user scenario describing a terminal task, paired with the model's reasoning chain (<think>) and a structured JSON answer (command + description). Dataset Summary Source model Qwen3.5-27B (DashScope API) Samples 3,632 Thinking mode Enabled (budget: 4096 tokens)… See the full description on the dataset page: https://huggingface.co/datasets/LocoreMind/qwen3.5-27b-cli-reasoning-3632x.texttext-generation1K<n<10K61 likes139 downloads7mo agoHugging Face25gavi56 /cli-1m CLI-1M: Industry-Diverse NL→Shell Training Corpus 975,933 natural-language → shell-command pairs · 18 industries · 6 shells · 13 languages · Apache-2.0 from datasets import load_dataset ds = load_dataset("carosh/cli-1m", revision="v1.0", split="train") # 843,461 rows — SFT-ready, license-filtered, quality-gated The most industry-diverse public dataset for NL→shell-command generation. 108× larger than NL2Bash (the previous public benchmark), and the first multilingual CLI… See the full description on the dataset page: https://huggingface.co/datasets/gavi56/cli-1m.texttext-generation1M<n<10M0 likes129 downloads15d agoHugging Face26mkurman /clinical-case-icd10-diagnosis Clinical History -> ICD-10 (Acute / Chronic) — CC BY-enriched 1798 de-identified clinical histories drawn from open-access case reports in PubMed Central, each paired with a single principal-diagnosis label: an ICD-10-CM code, its official descriptor, and an ACUTE/CHRONIC acuity status. input: a de-identified clinical history (presentation only; the diagnosis is removed and no PHI is present). output: {icd10_code, name, status} where status is ACUTE or CHRONIC (how the… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/clinical-case-icd10-diagnosis.texttext-classification1K<n<10K0 likes116 downloads3mo agoHugging Face27Akahsizrr /devin-cli-reasoning-distillation Devin CLI Reasoning Distillation Dataset A distillation dataset built from Devin CLI session traces, containing the model's internal reasoning traces (chain-of-thought / thinking), user prompts, assistant answers, and tool calls. The dataset is formatted to be directly compatible with SFT training pipelines that expect OpenAI-style message lists with a reasoning_content field. Dataset Summary Total rows 2,632 (2,507 train / 125 validation) Rows with… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/devin-cli-reasoning-distillation.tabulartext-generation1K<n<10K1 likes104 downloads12d agoHugging Face28JulesCan /clinicaltrial-protocol-corpus Clinical Trial Protocol Corpus Full-text clinical trial protocol documents from ClinicalTrials.gov with section segmentation aligned to SPIRIT/ICH-GCP categories. What this is 49,002 protocol PDFs downloaded from the ClinicalTrials.gov CDN, extracted to text via PyMuPDF, and segmented into structured sections. Each record contains the full protocol text plus a list of detected sections with headings, hierarchy levels, and section type labels from a 15-type… See the full description on the dataset page: https://huggingface.co/datasets/JulesCan/clinicaltrial-protocol-corpus.texttext-classification10K<n<100K0 likes93 downloads3mo agoHugging Face29Lots-of-LoRAs /task685_mmmlu_answer_generation_clinical_knowledge Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task685_mmmlu_answer_generation_clinical_knowledge Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task685_mmmlu_answer_generation_clinical_knowledge.texttext-generationn<1K0 likes89 downloads2y agoHugging Face30aisc-team-a1 /augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.texttext-generation10K<n<100K2 likes82 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.