CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kaptaan45 /KapInstruct-100M KapInstruct-100M: Curated 100-Million Token Instruction Tuning Dataset KapInstruct-100M is a high-fidelity, 100-million-token instruction-tuning dataset engineered for Supervised Fine-Tuning (SFT) and alignment of compact language models (under 1 billion parameters). Formatted with the Qwen ChatML chat template and tokenized using Qwen/Qwen3.5-0.8B-Base, the dataset enforces strict assistant-only loss masking (masking user prompts and structural delimiters to -100)… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapInstruct-100M.question-answering10K<n<100K0 likes631 downloads1mo agoHugging Face02lilywchen /lucky-initialization-atlas-100m-v2 Lucky initialization atlas v2 evidence Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs, provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses after those artifacts complete. It excludes credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B binary logs. tabulartext-generationn<1K0 likes296 downloads28d agoHugging Face03isthatshan /WestGenesis-Coder-SFT-100M Dataset Overview WestGenesis-Coder-Dataset is a meticulously curated coding dataset designed specifically for instruction-based model tuning and fine-tuning of existing models with enhanced code generation capabilities. This represents one of the largest and most comprehensively filtered corpora of publicly available coding data on the Hugging Face platform, with a non-thinking approach that emphasizes direct, concise code outputs for rapid model training. Key… See the full description on the dataset page: https://huggingface.co/datasets/isthatshan/WestGenesis-Coder-SFT-100M.texttext-generation1M<n<10M0 likes280 downloads3mo agoHugging Face04bbidpa /Rainbow-Pony-100m-Flutter-steps-eval Rainbow-Pony-100M Flutter — Steps Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-steps, a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files. In steps mode, the model is given an existing file and an edit instruction and generates a sequence of localized search/replace edit actions, each mechanically applied to the current file state before the next action is generated… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-steps-eval.tabulartext-generation1K<n<10K1 likes144 downloads17d agoHugging Face05bbidpa /Rainbow-Pony-100m-Flutter-direct-eval Rainbow-Pony-100M Flutter — Direct Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-direct, a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files. In direct mode, the model is given an existing file and an edit instruction and generates the complete modified file in a single forward pass (as opposed to the steps / iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-direct-eval.tabulartext-generation1K<n<10K0 likes102 downloads17d agoHugging Face06DimitarV /bulgarian-medical-cpt-100m Bulgarian text for MOSS continued pretraining Exactly 100 million training tokens: 10M medical and 90M general Bulgarian. An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers. No model training has been performed as part of this dataset build. Medical data Exactly 10,000,000 training tokens and 50,000 additional validation tokens, including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-100m.tabulartext-generation100K<n<1M0 likes88 downloads15d agoHugging Face07violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.tabulartext-generation1K<n<10K0 likes86 downloads2d agoHugging Face08violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.tabulartext-generation1K<n<10K0 likes85 downloads2d agoHugging Face09violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.tabulartext-generation1K<n<10K0 likes81 downloads2d agoHugging Face10Dodosoomro /simple-100m-pretrain-1b Simple-100M Pretraining Dataset (1B Tokens) A training-optimized, packed pretraining dataset for ~100M parameter language models. Built for reproducibility, minimal runtime overhead, and exact mixing ratios. 🎯 Purpose This dataset was created to train Simple-100M, a decoder-only Transformer targeting: ✅ Beat GPT-2-70M perplexity with minimal complexity ✅ Reproducible artifacts with exact token accounting ✅ Zero runtime preprocessing (ready-to-train) Target… See the full description on the dataset page: https://huggingface.co/datasets/Dodosoomro/simple-100m-pretrain-1b.texttext-generation1B<n<10B0 likes78 downloads6mo agoHugging Face11violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.tabulartext-generation1K<n<10K0 likes78 downloads2d agoHugging Face12codelion /sutra-100M Sutra 100M Pretraining Dataset A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.tabulartext-generation10K<n<100K3 likes61 downloads7mo agoHugging Face13krisbailey /RedPajama-Data-V2-100M RedPajama-Data-V2-100M Dataset Description This is a 100.0 Million token subset of krisbailey/RedPajama-Data-V2-1B, which is a subset of togethercomputer/RedPajama-Data-V2. Motivation 100M tokens is a standard size for: CI/CD Pipelines: Fast enough to download and train for unit tests. Debugging: Verifying training loops without waiting for hours. Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B). Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-100M.texttext-generation10K<n<100K0 likes58 downloads8mo agoHugging Face14hanspeterlyngsoeraaschoujensen /swerebench-openhands-100m-max64k SWE-rebench OpenHands 100M SFT Subset (max 64k) This is a deterministic, representative subset of nebius/SWE-rebench-openhands-trajectories, augmented with exact sequence and supervised-loss token counts. The source trajectories were collected with Qwen3-Coder-480B-A35B-Instruct and OpenHands v0.54.0. This derivative preserves the source dataset's CC BY 4.0 license and attribution. Filters and size 7,867 trajectories 100,095,655 assistant loss tokens 352,709,237… See the full description on the dataset page: https://huggingface.co/datasets/hanspeterlyngsoeraaschoujensen/swerebench-openhands-100m-max64k.tabulartext-generation1K<n<10K0 likes57 downloads2mo agoHugging Face15nielsr /MS-GPT-Pretraining-100M MS-GPT 100M pretraining corpus This repository contains the molecule-only pretraining corpus released with MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model. The released metadata reports 100,007,359 molecule records, 4,096 fingerprint bits, and 9,413 excluded InChIKeys. The corpus is accompanied by formula vectors and formula-group metadata. File Description data.arrow Molecule pretraining… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/MS-GPT-Pretraining-100M.text-generation100M<n<1B0 likes56 downloads2mo agoHugging Face16violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think 4,000 evaluation attempts: 250 tasks × 16 samples. Original samples 0–3 and twelve additional seeded samples 4–15. The train split contains evaluation records, not training data. The evaluated model is violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m at revision 0c295885100d6eba4f514752aa081c5b0c73fdec. Cohort Attempts All-criteria-pass rate ± task-level SEM Original four 1,000 8.000% ±… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think.tabulartext-generation1K<n<10K0 likes54 downloads8h agoHugging Face17codelion /sutra-improved-100M Sutra Improved 100M A self-improved pedagogical dataset for LLM pretraining, containing 413,899 entries totaling 110,038,011 tokens (~110 million). This dataset was created by applying an iterative self-improvement process to the Sutra-10B dataset, where each sample was rewritten using Gemma-3-4B-IT and only the better version (original or rewritten) was kept, followed by comprehensive deduplication and quality filtering. Dataset Description This dataset explores… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-improved-100M.texttext-generation100K<n<1M2 likes39 downloads6mo agoHugging Face18krisbailey /cosmopedia-100M cosmopedia-100M Dataset Description This is a 100.0 Million token subset of krisbailey/cosmopedia-1B, which is a subset of HuggingFaceTB/cosmopedia. Motivation 100M tokens is a standard size for: CI/CD Pipelines: Fast enough to download and train for unit tests. Debugging: Verifying training loops without waiting for hours. Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B). Dataset Details Total Tokens: 100,000,060… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-100M.texttext-generation100K<n<1M0 likes33 downloads8mo agoHugging Face19CausalNLP /clt-tokenized-control-ar-zh-ko-ja-100m Balanced Arabic-Chinese-Korean-Japanese 4B-token training data Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 100,000,000 tokens in complete documents. Total target: 400,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard. Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>. texttext-generation100K<n<1M0 likes31 downloads3mo agoHugging Face20VertexResearch /Vertex-0.6-100M-self-identification Vertex 0.6 100M Self Identification A self-identification SFT dataset for Vertex-0.6-100M-8192-Instruct: 459 ChatML-style conversations that teach the model who it is: its name, creator, family, architecture, parameter count and knowledge cutoff. Made from SupraLabs/LLM-self-identification (Apache-2.0), with every {{SELF_ID.*}} marker replaced with the Vertex 0.6 100M identity: Marker Value MODEL_ID VertexResearch/Vertex-0.6-100M-8192-Instruct MODEL_NAME Vertex 0.6… See the full description on the dataset page: https://huggingface.co/datasets/VertexResearch/Vertex-0.6-100M-self-identification.texttext-generationn<1K0 likes20 downloads2d agoHugging Face21zhiwei555 /dclm_data_100m Data-Constrained Language Model Pretraining (100M) This dataset contains pre-tokenized .pt files containing packed GPT-2-tokenized sequences. It was used in the research presented in the paper Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws. The repository provides the 100M token training split along with a validation split, specifically prepared for experiments in data-constrained, compute-rich regimes. Links Paper:… See the full description on the dataset page: https://huggingface.co/datasets/zhiwei555/dclm_data_100m.text-generation1 likes16 downloads4mo agoHugging Face22blab-jhu /KYS-DCLM-Refinedweb-100M-Scoredgated KYS-DCLM-Refinedweb-100M-Scored The candidate document pool behind Know Your Sources: Data Selection Matters when Rewriting for Data-Constrained Pretraining — 99,949,162 web documents sampled from DCLM-RefinedWeb, each annotated with three independent quality scores, their tie-aware global percentiles, a 24-way WebOrganizer topic label, and a Llama-2 token count. Every source-selection strategy in the paper is a different way of ranking this table. Contents 200… See the full description on the dataset page: https://huggingface.co/datasets/blab-jhu/KYS-DCLM-Refinedweb-100M-Scored.tabulartext-generation10M<n<100M0 likes13 downloads1mo agoHugging Face23krisbailey /falcon-refinedweb-100M falcon-refinedweb-100M Dataset Description This is a 100.0 Million token subset of krisbailey/falcon-refinedweb-1B, which is a subset of tiiuae/falcon-refinedweb. Motivation 100M tokens is a standard size for: CI/CD Pipelines: Fast enough to download and train for unit tests. Debugging: Verifying training loops without waiting for hours. Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B). Dataset Details Total Tokens:… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-100M.texttext-generation100K<n<1M0 likes10 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.