CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KantaHayashiAI /ClimbLab-JaJapanese / 日本語版 ClimbLab-Ja ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.tabulartext-generation100M<n<1B2 likes10k downloads4mo agoHugging Face02OptimalScale /ClimbLabClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters. Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.texttext-generation1B<n<10B16 likes9.8k downloads1y agoHugging Face03Yujivus /nanochat-climbmix-arithmetic-base10 nanochat ClimbMix + Base-10 Arithmetic This dataset contains the first 170 shuffled ClimbMix training shards used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is mixed into shards 00000..00149; the final 20 train shards are unchanged web-only padding. The original validation shard (shard_06542.parquet) is also copied unchanged. Arithmetic corpus Family Examples a + b = c (all ordered pairs 0..2000, two exposures) 8,008,002 a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.texttext-generation10M<n<100M0 likes3.8k downloads1mo agoHugging Face04OptimalScale /ClimbMixClimbMix is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbMix is a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. It was introduced in this paper. We proposed a new algorithm to filter and mix the dataset. First, we grouped the data into 1,000 groups based on topic information. Then we applied two classifiers: one to detect advertisements and another to… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbMix.tabulartext-generation100M<n<1B36 likes3.2k downloads1y agoHugging Face05Yujivus /nanochat-climbmix-arithmetic-base7 nanochat ClimbMix + Arithmetic: base-7 numeral world This is a deterministic base-7 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 7. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base7.texttext-generation10M<n<100M0 likes1.5k downloads1mo agoHugging Face06AGBonnet /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.texttext-generation10K<n<100K75 likes1.3k downloads3y agoHugging Face07Yujivus /nanochat-climbmix-170 nanochat ClimbMix: first 170 train shards Convenience mirror of the exact initial ClimbMix slice downloaded by python -m nanochat.dataset -n 170. Contents Training: shard_00000.parquet through shard_00169.parquet Validation: shard_06542.parquet manifest.json: pinned source revision, file list, and byte sizes The Parquet shards are copied without modifying their rows or text. Attribution and provenance nanochat:… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-170.texttext-generation10M<n<100M0 likes928 downloads1mo agoHugging Face08starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes918 downloads2y agoHugging Face09LocalDoc /climbmix-40b-az ClimbMix 40B — Azerbaijani A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate. Dataset Summary This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language. Property Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.texttext-generation10M<n<100M1 likes647 downloads6mo agoHugging Face10Yujivus /nanochat-climbmix-arithmetic-base6 nanochat ClimbMix + Arithmetic: base-6 numeral world This is a deterministic base-6 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 6. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base6.texttext-generation10M<n<100M0 likes579 downloads1mo agoHugging Face11Clinton /Text-to-sql-v1texttext-generation100K<n<1M73 likes343 downloads3y agoHugging Face12stindardlogic /medical-clinical-reasoning-sft-100k Medical Clinical Reasoning SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education. Dataset Description This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.texttext-generation100K<n<1M0 likes333 downloads2mo agoHugging Face13hugo /protocolos-clinicos-br Protocolos Clínicos BR Paper | Code | Blog post Brazilian Ministry of Health official clinical guidelines (PCDTs and related) plus the synthetic training corpus derived from them, used to adapt LLMs to Brazilian clinical knowledge. This dataset was introduced in the paper "Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines". Configurations default — Original guidelines (raw text) The 178 official Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/hugo/protocolos-clinicos-br.texttext-generation10K<n<100K0 likes317 downloads2mo agoHugging Face14carosh /cli-1m CLI-1M: Industry-Diverse NL→Shell Training Corpus 975,933 natural-language → shell-command pairs · 18 industries · 6 shells · 13 languages · Apache-2.0 from datasets import load_dataset ds = load_dataset("carosh/cli-1m", revision="v1.0", split="train") # 843,461 rows — SFT-ready, license-filtered, quality-gated The most industry-diverse public dataset for NL→shell-command generation. 108× larger than NL2Bash (the previous public benchmark), and the first multilingual CLI corpus.… See the full description on the dataset page: https://huggingface.co/datasets/carosh/cli-1m.texttext-generation1M<n<10M1 likes237 downloads4mo agoHugging Face15Giordanopsouza /clinicalbr ClinicalBr ClinicalBr is the first bilingual (Portuguese–English) clinical-decision benchmark built from 2,892 real Brazilian case reports drawn from 36 open-access medical journals spanning 18 specialties. Every case is provided as a parallel PT/EN pair and supports four evaluation tasks. Please refer to the paper for full details on the tasks, methodology, and limitations. Tasks & metrics Task Config n / lang Metric Diagnosis retrieval diagnosis 2,135… See the full description on the dataset page: https://huggingface.co/datasets/Giordanopsouza/clinicalbr.textquestion-answering10K<n<100K0 likes156 downloads1mo agoHugging Face16b-mc2 /cli-commands-explained Overview This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.tabulartext-generation10K<n<100K5 likes151 downloads2y agoHugging Face17C-lister /ChainSWE ChainSWE ChainSWE is a benchmark of sequential, dependent bug fixes for evaluating coding agents on continuous software maintenance. It contains 100 time-ordered chains (304 bug-fix tasks) mined from six SWE-bench-family datasets across 54 Python repositories; each row is one chain over a single repository sharing one base commit and pre-built Docker image, and its bug_fixes field lists the tasks in chronological order, where each task is a self-contained SWE-bench-style… See the full description on the dataset page: https://huggingface.co/datasets/C-lister/ChainSWE.texttext-generationn<1K0 likes150 downloads3mo agoHugging Face18LocoreMind /qwen3.5-27b-cli-reasoning-3632x Qwen3.5-27B CLI Reasoning 3632x A synthetic reasoning dataset for CLI/terminal command assistance, distilled from Qwen3.5-27B with thinking mode enabled. Each sample contains a realistic user scenario describing a terminal task, paired with the model's reasoning chain (<think>) and a structured JSON answer (command + description). Dataset Summary Source model Qwen3.5-27B (DashScope API) Samples 3,632 Thinking mode Enabled (budget: 4096 tokens)… See the full description on the dataset page: https://huggingface.co/datasets/LocoreMind/qwen3.5-27b-cli-reasoning-3632x.texttext-generation1K<n<10K61 likes136 downloads7mo agoHugging Face19gavi56 /cli-1m CLI-1M: Industry-Diverse NL→Shell Training Corpus 975,933 natural-language → shell-command pairs · 18 industries · 6 shells · 13 languages · Apache-2.0 from datasets import load_dataset ds = load_dataset("carosh/cli-1m", revision="v1.0", split="train") # 843,461 rows — SFT-ready, license-filtered, quality-gated The most industry-diverse public dataset for NL→shell-command generation. 108× larger than NL2Bash (the previous public benchmark), and the first multilingual CLI… See the full description on the dataset page: https://huggingface.co/datasets/gavi56/cli-1m.texttext-generation1M<n<10M0 likes133 downloads16d agoHugging Face20mkurman /clinical-case-icd10-diagnosis Clinical History -> ICD-10 (Acute / Chronic) — CC BY-enriched 1798 de-identified clinical histories drawn from open-access case reports in PubMed Central, each paired with a single principal-diagnosis label: an ICD-10-CM code, its official descriptor, and an ACUTE/CHRONIC acuity status. input: a de-identified clinical history (presentation only; the diagnosis is removed and no PHI is present). output: {icd10_code, name, status} where status is ACUTE or CHRONIC (how the… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/clinical-case-icd10-diagnosis.texttext-classification1K<n<10K0 likes119 downloads3mo agoHugging Face21Akahsizrr /devin-cli-reasoning-distillation Devin CLI Reasoning Distillation Dataset A distillation dataset built from Devin CLI session traces, containing the model's internal reasoning traces (chain-of-thought / thinking), user prompts, assistant answers, and tool calls. The dataset is formatted to be directly compatible with SFT training pipelines that expect OpenAI-style message lists with a reasoning_content field. Dataset Summary Total rows 2,632 (2,507 train / 125 validation) Rows with… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/devin-cli-reasoning-distillation.tabulartext-generation1K<n<10K1 likes106 downloads13d agoHugging Face22JulesCan /clinicaltrial-protocol-corpus Clinical Trial Protocol Corpus Full-text clinical trial protocol documents from ClinicalTrials.gov with section segmentation aligned to SPIRIT/ICH-GCP categories. What this is 49,002 protocol PDFs downloaded from the ClinicalTrials.gov CDN, extracted to text via PyMuPDF, and segmented into structured sections. Each record contains the full protocol text plus a list of detected sections with headings, hierarchy levels, and section type labels from a 15-type… See the full description on the dataset page: https://huggingface.co/datasets/JulesCan/clinicaltrial-protocol-corpus.texttext-classification10K<n<100K0 likes94 downloads4mo agoHugging Face23Lots-of-LoRAs /task685_mmmlu_answer_generation_clinical_knowledge Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task685_mmmlu_answer_generation_clinical_knowledge Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task685_mmmlu_answer_generation_clinical_knowledge.texttext-generationn<1K0 likes91 downloads2y agoHugging Face24aisc-team-a1 /augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.texttext-generation10K<n<100K2 likes85 downloads3y agoHugging Face25johnny8808 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/johnny8808/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes76 downloads5mo agoHugging Face26AmareshHebbar /clinical-summarizer-sft Clinical Note Summarizer (SOAP Format) Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Long clinical notes → structured SOAP summaries Why download this Automate clinical documentation. Reduce physician burnout by summarizing visit notes into Subjective / Objective / Assessment / Plan format. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/clinical-summarizer-sft.texttext-generation10K<n<100K0 likes74 downloads3mo agoHugging Face27shariqazeem /psiddx-clinical-ddx PsiDDx — Clinical Differential-Diagnosis Dataset (v5.2) A decontaminated, calibrated differential-diagnosis training corpus for the Adaption Labs AutoScientist Challenge (Healthcare). Each record pairs a patient presentation with a ranked top-5 differential — common explanations first, each with an ICD-10 code, a calibrated confidence, and the discriminating feature; rare diagnoses are kept as low-confidence must-not-miss entries rather than over-called. 5,482 rows. Used to… See the full description on the dataset page: https://huggingface.co/datasets/shariqazeem/psiddx-clinical-ddx.texttext-generation1K<n<10K0 likes71 downloads3mo agoHugging Face28cyrilzakka /clinical-trials-embeddings Clinical Trials Embeddings Dataset Overview This dataset contains information extracted from clinical trial records collected from ClinicalTrials.gov (Date Accessed: 05/02/2025) along with briefSummary columns embeddings generated using minishlab/potion-base-8M. It focuses on key descriptive fields that provide insight into trial objectives, eligibility criteria, and study design. The dataset is designed for researchers, healthcare professionals, and AI/ML practitioners… See the full description on the dataset page: https://huggingface.co/datasets/cyrilzakka/clinical-trials-embeddings.texttext-classification100K<n<1M4 likes67 downloads1y agoHugging Face29TonicAI /synthetic_clinical_conversations Synthetic Clinical Conversations Fully synthetic English clinical conversations (care-coordination calls, telehealth visits, post-discharge check-ins) paired with structured encounter records, built for training and evaluating transcript→JSON extraction models. Generated structure-first with Tonic Fabricate: the structured facts are authored as relational data with controlled vocabularies, the conversation is rendered from those facts, and the extraction target is a… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/synthetic_clinical_conversations.texttext-generation1K<n<10K0 likes67 downloads2mo agoHugging Face30Vinay393 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Vinay393/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes66 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.