CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lilywchen /lucky-initialization-atlas-100m-v2 Lucky initialization atlas v2 evidence Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs, provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses after those artifacts complete. It excludes credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B binary logs. tabulartext-generationn<1K0 likes296 downloads28d agoHugging Face02bbidpa /Rainbow-Pony-100m-Flutter-steps-eval Rainbow-Pony-100M Flutter — Steps Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-steps, a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files. In steps mode, the model is given an existing file and an edit instruction and generates a sequence of localized search/replace edit actions, each mechanically applied to the current file state before the next action is generated… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-steps-eval.tabulartext-generation1K<n<10K1 likes144 downloads17d agoHugging Face03bbidpa /Rainbow-Pony-100m-Flutter-direct-eval Rainbow-Pony-100M Flutter — Direct Mode — Validation Results Dataset Summary Held-out evaluation results for bbidpa/Rainbow-Pony-100m-Flutter-direct, a 100M-parameter transformer trained from scratch to edit Flutter/Dart source files. In direct mode, the model is given an existing file and an edit instruction and generates the complete modified file in a single forward pass (as opposed to the steps / iterative diff-based mode — see the sibling dataset… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/Rainbow-Pony-100m-Flutter-direct-eval.tabulartext-generation1K<n<10K0 likes102 downloads17d agoHugging Face04DimitarV /bulgarian-medical-cpt-100m Bulgarian text for MOSS continued pretraining Exactly 100 million training tokens: 10M medical and 90M general Bulgarian. An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers. No model training has been performed as part of this dataset build. Medical data Exactly 10,000,000 training tokens and 50,000 additional validation tokens, including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-100m.tabulartext-generation100K<n<1M0 likes88 downloads15d agoHugging Face05violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.tabulartext-generation1K<n<10K0 likes86 downloads2d agoHugging Face06violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.tabulartext-generation1K<n<10K0 likes85 downloads2d agoHugging Face07violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.tabulartext-generation1K<n<10K0 likes81 downloads2d agoHugging Face08violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.tabulartext-generation1K<n<10K0 likes78 downloads2d agoHugging Face09codelion /sutra-100M Sutra 100M Pretraining Dataset A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.tabulartext-generation10K<n<100K3 likes61 downloads7mo agoHugging Face10hanspeterlyngsoeraaschoujensen /swerebench-openhands-100m-max64k SWE-rebench OpenHands 100M SFT Subset (max 64k) This is a deterministic, representative subset of nebius/SWE-rebench-openhands-trajectories, augmented with exact sequence and supervised-loss token counts. The source trajectories were collected with Qwen3-Coder-480B-A35B-Instruct and OpenHands v0.54.0. This derivative preserves the source dataset's CC BY 4.0 license and attribution. Filters and size 7,867 trajectories 100,095,655 assistant loss tokens 352,709,237… See the full description on the dataset page: https://huggingface.co/datasets/hanspeterlyngsoeraaschoujensen/swerebench-openhands-100m-max64k.tabulartext-generation1K<n<10K0 likes57 downloads2mo agoHugging Face11violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think 4,000 evaluation attempts: 250 tasks × 16 samples. Original samples 0–3 and twelve additional seeded samples 4–15. The train split contains evaluation records, not training data. The evaluated model is violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m at revision 0c295885100d6eba4f514752aa081c5b0c73fdec. Cohort Attempts All-criteria-pass rate ± task-level SEM Original four 1,000 8.000% ±… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think.tabulartext-generation1K<n<10K0 likes54 downloads2h agoHugging Face12blab-jhu /KYS-DCLM-Refinedweb-100M-Scoredgated KYS-DCLM-Refinedweb-100M-Scored The candidate document pool behind Know Your Sources: Data Selection Matters when Rewriting for Data-Constrained Pretraining — 99,949,162 web documents sampled from DCLM-RefinedWeb, each annotated with three independent quality scores, their tie-aware global percentiles, a 24-way WebOrganizer topic label, and a Llama-2 token count. Every source-selection strategy in the paper is a different way of ranking this table. Contents 200… See the full description on the dataset page: https://huggingface.co/datasets/blab-jhu/KYS-DCLM-Refinedweb-100M-Scored.tabulartext-generation10M<n<100M0 likes13 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.