CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lilywchen /lucky-initialization-atlas-100m-v2 Lucky initialization atlas v2 evidence Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs, provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses after those artifacts complete. It excludes credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B binary logs. tabulartext-generationn<1K0 likes296 downloads28d agoHugging Face02isthatshan /WestGenesis-Coder-SFT-100M Dataset Overview WestGenesis-Coder-Dataset is a meticulously curated coding dataset designed specifically for instruction-based model tuning and fine-tuning of existing models with enhanced code generation capabilities. This represents one of the largest and most comprehensively filtered corpora of publicly available coding data on the Hugging Face platform, with a non-thinking approach that emphasizes direct, concise code outputs for rapid model training. Key… See the full description on the dataset page: https://huggingface.co/datasets/isthatshan/WestGenesis-Coder-SFT-100M.texttext-generation1M<n<10M0 likes280 downloads3mo agoHugging Face03bbidpa /Rainbow-Pony-100m-Flutter-steps-eval Rainbow-Pony-100M Flutter — Steps Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-steps, a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files. In steps mode, the model is given an existing file and an edit instruction and generates a sequence of localized search/replace edit actions, each mechanically applied to the current file state before the next action is generated… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-steps-eval.tabulartext-generation1K<n<10K1 likes144 downloads17d agoHugging Face04bbidpa /Rainbow-Pony-100m-Flutter-direct-eval Rainbow-Pony-100M Flutter — Direct Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-direct, a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files. In direct mode, the model is given an existing file and an edit instruction and generates the complete modified file in a single forward pass (as opposed to the steps / iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-direct-eval.tabulartext-generation1K<n<10K0 likes102 downloads17d agoHugging Face05DimitarV /bulgarian-medical-cpt-100m Bulgarian text for MOSS continued pretraining Exactly 100 million training tokens: 10M medical and 90M general Bulgarian. An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers. No model training has been performed as part of this dataset build. Medical data Exactly 10,000,000 training tokens and 50,000 additional validation tokens, including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-100m.tabulartext-generation100K<n<1M0 likes88 downloads14d agoHugging Face06violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.tabulartext-generation1K<n<10K0 likes86 downloads2d agoHugging Face07violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.tabulartext-generation1K<n<10K0 likes85 downloads2d agoHugging Face08violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.tabulartext-generation1K<n<10K0 likes81 downloads2d agoHugging Face09Dodosoomro /simple-100m-pretrain-1b Simple-100M Pretraining Dataset (1B Tokens) A training-optimized, packed pretraining dataset for ~100M parameter language models. Built for reproducibility, minimal runtime overhead, and exact mixing ratios. 🎯 Purpose This dataset was created to train Simple-100M, a decoder-only Transformer targeting: ✅ Beat GPT-2-70M perplexity with minimal complexity ✅ Reproducible artifacts with exact token accounting ✅ Zero runtime preprocessing (ready-to-train) Target… See the full description on the dataset page: https://huggingface.co/datasets/Dodosoomro/simple-100m-pretrain-1b.texttext-generation1B<n<10B0 likes78 downloads6mo agoHugging Face10violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.tabulartext-generation1K<n<10K0 likes78 downloads2d agoHugging Face11codelion /sutra-100M Sutra 100M Pretraining Dataset A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.tabulartext-generation10K<n<100K3 likes61 downloads7mo agoHugging Face12krisbailey /RedPajama-Data-V2-100M RedPajama-Data-V2-100M Dataset Description This is a 100.0 Million token subset of krisbailey/RedPajama-Data-V2-1B, which is a subset of togethercomputer/RedPajama-Data-V2. Motivation 100M tokens is a standard size for: CI/CD Pipelines: Fast enough to download and train for unit tests. Debugging: Verifying training loops without waiting for hours. Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B). Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-100M.texttext-generation10K<n<100K0 likes58 downloads8mo agoHugging Face13hanspeterlyngsoeraaschoujensen /swerebench-openhands-100m-max64k SWE-rebench OpenHands 100M SFT Subset (max 64k) This is a deterministic, representative subset of nebius/SWE-rebench-openhands-trajectories, augmented with exact sequence and supervised-loss token counts. The source trajectories were collected with Qwen3-Coder-480B-A35B-Instruct and OpenHands v0.54.0. This derivative preserves the source dataset's CC BY 4.0 license and attribution. Filters and size 7,867 trajectories 100,095,655 assistant loss tokens 352,709,237… See the full description on the dataset page: https://huggingface.co/datasets/hanspeterlyngsoeraaschoujensen/swerebench-openhands-100m-max64k.tabulartext-generation1K<n<10K0 likes57 downloads2mo agoHugging Face14violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and binary all-criteria-pass rule. Mean all-pass rate: 8.0000%. The train split contains held-out evaluation records, not training examples. Model and training mixture The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think.tabulartext-generation1K<n<10K0 likes54 downloads4d agoHugging Face15codelion /sutra-improved-100M Sutra Improved 100M A self-improved pedagogical dataset for LLM pretraining, containing 413,899 entries totaling 110,038,011 tokens (~110 million). This dataset was created by applying an iterative self-improvement process to the Sutra-10B dataset, where each sample was rewritten using Gemma-3-4B-IT and only the better version (original or rewritten) was kept, followed by comprehensive deduplication and quality filtering. Dataset Description This dataset explores… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-improved-100M.texttext-generation100K<n<1M2 likes39 downloads6mo agoHugging Face16krisbailey /cosmopedia-100M cosmopedia-100M Dataset Description This is a 100.0 Million token subset of krisbailey/cosmopedia-1B, which is a subset of HuggingFaceTB/cosmopedia. Motivation 100M tokens is a standard size for: CI/CD Pipelines: Fast enough to download and train for unit tests. Debugging: Verifying training loops without waiting for hours. Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B). Dataset Details Total Tokens: 100,000,060… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-100M.texttext-generation100K<n<1M0 likes33 downloads8mo agoHugging Face17CausalNLP /clt-tokenized-control-ar-zh-ko-ja-100m Balanced Arabic-Chinese-Korean-Japanese 4B-token training data Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 100,000,000 tokens in complete documents. Total target: 400,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard. Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>. texttext-generation100K<n<1M0 likes31 downloads3mo agoHugging Face18VertexResearch /Vertex-0.6-100M-self-identification Vertex 0.6 100M Self Identification A self-identification SFT dataset for Vertex-0.6-100M-8192-Instruct: 459 ChatML-style conversations that teach the model who it is: its name, creator, family, architecture, parameter count and knowledge cutoff. Made from SupraLabs/LLM-self-identification (Apache-2.0), with every {{SELF_ID.*}} marker replaced with the Vertex 0.6 100M identity: Marker Value MODEL_ID VertexResearch/Vertex-0.6-100M-8192-Instruct MODEL_NAME Vertex 0.6… See the full description on the dataset page: https://huggingface.co/datasets/VertexResearch/Vertex-0.6-100M-self-identification.texttext-generationn<1K0 likes20 downloads2d agoHugging Face19blab-jhu /KYS-DCLM-Refinedweb-100M-Scoredgated KYS-DCLM-Refinedweb-100M-Scored The candidate document pool behind Know Your Sources: Data Selection Matters when Rewriting for Data-Constrained Pretraining — 99,949,162 web documents sampled from DCLM-RefinedWeb, each annotated with three independent quality scores, their tie-aware global percentiles, a 24-way WebOrganizer topic label, and a Llama-2 token count. Every source-selection strategy in the paper is a different way of ranking this table. Contents 200… See the full description on the dataset page: https://huggingface.co/datasets/blab-jhu/KYS-DCLM-Refinedweb-100M-Scored.tabulartext-generation10M<n<100M0 likes13 downloads1mo agoHugging Face20krisbailey /falcon-refinedweb-100M falcon-refinedweb-100M Dataset Description This is a 100.0 Million token subset of krisbailey/falcon-refinedweb-1B, which is a subset of tiiuae/falcon-refinedweb. Motivation 100M tokens is a standard size for: CI/CD Pipelines: Fast enough to download and train for unit tests. Debugging: Verifying training loops without waiting for hours. Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B). Dataset Details Total Tokens:… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-100M.texttext-generation100K<n<1M0 likes10 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.