CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SlayerLab /minimal-en-corpus-5b Minimal EN Corpus 5B An English-language pretraining corpus prepared for controlled experiments with approximately 125M-parameter GPT-2 models based on karpathy/nanoGPT. The name refers to the approximately 5B-token mixture selected before final BPE tokenization. With the included 12,288-token BPE tokenizer, the packaged nanoGPT training split contains 5,396,605,407 tokens. Contents The dataset provides both reusable source text and ready-to-train nanoGPT… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b.text1M<n<10M1 likes2.3k downloads2d agoHugging Face02geodesic-research /pa-warm-start-sft-medium-5b-mix geodesic-research/pa-warm-start-sft-medium-5b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-medium-5b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-medium-5b-mix.tabular1M<n<10M0 likes799 downloads1mo agoHugging Face03idkdd /icrm-hitek-full-db-mixed-5b ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet). Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent). Contents The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/idkdd/icrm-hitek-full-db-mixed-5b.text1B<n<10B0 likes582 downloads15d agoHugging Face04PatrickHaller /fineweb-5Btext1M<n<10M0 likes452 downloads2y agoHugging Face05skymizer /fineweb-edu-dedup-5Btext1M<n<10M0 likes408 downloads2y agoHugging Face06sfanm /d24-midtrain-olmo3-5b d24 Midtrain — OLMo-3 Dolmino (5B, chunked) A 5.00B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix, built by taking a fraction of each component of allenai/dolma3_dolmino_mix-100B-1025. Unlike the smaller d24-midtrain-olmo3 (which dropped documents over 2048 tokens — silently zeroing the long reasoning-trace and PDF components), this build chunks long documents into 2048-token windows (decoded back to text), so every component keeps ~100% of its tokens and reaches… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b.texttext-generation10M<n<100M0 likes346 downloads3mo agoHugging Face07openbmb /InfLLM-V2-data-5B InfLLM-V2 Long-Context Training Dataset with 5B Tokens Project Links: [Paper] [InfLLM-V2 Models] [CUDA Kernel Code] 🚀 About InfLLM-V2 InfLLM-V2 is a native sparse attention framework designed for the efficient processing of long-sequence texts. Its core advantage is the ability to maintain high performance comparable to dense attention in short-text scenarios—without any extra parameters—while seamlessly switching to a sparse mode for long-text scenarios, achieving… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/InfLLM-V2-data-5B.text1M<n<10M36 likes316 downloads11mo agoHugging Face08sfanm /d24-midtrain-olmo3-5b-wholedoc d24 Midtrain — OLMo-3 Dolmino (5B, whole-doc) A 4.71B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix at its exact component proportions, built by taking a fraction of each component of allenai/dolma3_dolmino_mix-100B-1025. Documents are kept whole — no length filter, no chunking. Why whole-doc: for packed pretrain/midtrain the trainer (e.g. Megatron preprocess_data --append-eod GPTDataset) already concatenates documents and slices them into context-length windows… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b-wholedoc.texttext-generation10M<n<100M0 likes292 downloads3mo agoHugging Face09electricsheepafrica /africa-sudan-sudan-environment-5b08204c Sudan - Environment | Africa (Sudan official open data) 4,043 rows - 1 Africa country - 1961-2025 - Repackaged by Electric Sheep Africa TL;DR This dataset packages one official CSV resource from Sudan as ML-ready Parquet. The source file is the provenance boundary; all usable indicators or tabular columns from the resource stay together in this repo. About the source Source: Sudan - Environment Publisher: World Bank Group Resource: Environment… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-sudan-sudan-environment-5b08204c.tabulartabular-regression1K<n<10K0 likes272 downloads1mo agoHugging Face10textcleanlm /med-domain-5btext1M<n<10M0 likes241 downloads1y agoHugging Face11luckysantoso /rmfs-wm-v3-jenkins5b RMFS World Model v3 — jenkins5-b1 Egocentric driving data for a single robot (the ego) in a simulated Robotic Mobile Fulfillment System (RMFS) warehouse, collected from RAWSim-O via a custom Gym server. Intended for training World Models-style V (VAE) + M (MDN-RNN) + C (controller) stacks where the controller is trained entirely inside the learned dream. At a glance Episodes 120 Frames 512,734 Hops (macro-steps) 28,800 Cameras 4 x 96x96 RGB… See the full description on the dataset page: https://huggingface.co/datasets/luckysantoso/rmfs-wm-v3-jenkins5b.tabularroboticsn<1K0 likes214 downloads27d agoHugging Face12lparkourer10 /starcoder-python5b5b gpt2 tokens tabulartext-generation1M<n<10M0 likes205 downloads2y agoHugging Face13happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_reg_grpo_bce_4textn<1K0 likes204 downloads1y agoHugging Face14synpre /dclm_seed_5b_tanishqtabular1M<n<10M0 likes192 downloads2y agoHugging Face15elonmuskceo /InfLLM-V2-data-5B-v2 InfLLM-V2 Long-Context Training Dataset with 5B Tokens Project Links: [Paper] [InfLLM-V2 Models] [CUDA Kernel Code] 🚀 About InfLLM-V2 InfLLM-V2 is a native sparse attention framework designed for the efficient processing of long-sequence texts. Its core advantage is the ability to maintain high performance comparable to dense attention in short-text scenarios—without any extra parameters—while seamlessly switching to a sparse mode for long-text scenarios, achieving… See the full description on the dataset page: https://huggingface.co/datasets/elonmuskceo/InfLLM-V2-data-5B-v2.text1M<n<10M0 likes188 downloads10mo agoHugging Face16happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_4textn<1K0 likes182 downloads1y agoHugging Face17happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_2textn<1K0 likes162 downloads1y agoHugging Face18happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gpg_bce_5textn<1K0 likes159 downloads1y agoHugging Face19happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_5textn<1K0 likes148 downloads1y agoHugging Face20cudabenchmarktest /r8-eval-suite-5bucket ⚠️ CRITICAL: Ollama Inference Flag Required for derived models If you train or serve any Qwen3.5-9B-derived model from this lineage via Ollama, you MUST pass "think": false in /api/chat requests for chat / instruction following / tool use. The qwen3.5 RENDERER auto-injects <think> tags causing 25-46% empty-answer rates without this flag. See dataset cudabenchmarktest/r9-research-framework/_OLLAMA_INFERENCE_WARNING.md for the full lesson learned. R8/R9 Five-Bucket… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r8-eval-suite-5bucket.tabulartext-generationn<1K0 likes131 downloads5mo agoHugging Face21elonmuskceo /InfLLM-V2-data-5B InfLLM-V2 Long-Context Training Dataset with 5B Tokens Project Links: [Paper] [InfLLM-V2 Models] [CUDA Kernel Code] 🚀 About InfLLM-V2 InfLLM-V2 is a native sparse attention framework designed for the efficient processing of long-sequence texts. Its core advantage is the ability to maintain high performance comparable to dense attention in short-text scenarios—without any extra parameters—while seamlessly switching to a sparse mode for long-text scenarios, achieving… See the full description on the dataset page: https://huggingface.co/datasets/elonmuskceo/InfLLM-V2-data-5B.text1M<n<10M0 likes127 downloads10mo agoHugging Face22BroAlanTaps /GPT2-Large-Llama3-8B-fineweb-256-5Btokens Source: Modified from HuggingFaceFW/fineweb Sample-10BT subset. The preprocessed script files are in the data directory. You can download it and run: python fineweb.py --segment_length 512 --nproc 32 --batch_size 1024 --save_path ./fineweb10B/save/ Dataset Owner(s): Individual: BroAlanTaps License/Terms of Use: odc-by: Open Data Commons License Attribution family Intended Usage: This dataset is intended to be used with research or… See the full description on the dataset page: https://huggingface.co/datasets/BroAlanTaps/GPT2-Large-Llama3-8B-fineweb-256-5Btokens.text10M<n<100M1 likes119 downloads1y agoHugging Face23icedpanda /bright-passage-index-gte_qwen2-1_5btext1M<n<10M0 likes106 downloads1y agoHugging Face24happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gspo_bce_4textn<1K0 likes98 downloads1y agoHugging Face25jprivera44 /atlas9_5beh_sequential_sdftext100K<n<1M0 likes83 downloads16d agoHugging Face26happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gspo_bce_3textn<1K0 likes78 downloads1y agoHugging Face27ssuresh /codelion_all-5Btext1M<n<10M0 likes65 downloads5mo agoHugging Face28happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gspo_bce_5textn<1K0 likes62 downloads1y agoHugging Face29happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_reg_grpo_bce_2textn<1K0 likes60 downloads1y agoHugging Face30electricsheepafrica /africa-cote-d-ivoire-sites-de-vaccination-de-covid-19-dans-le-district-d-abidja-5bbc9c04 Sites De Vaccination De Covid 19 Dans Le District D Abidja | Africa (Cote d'Ivoire DataFair) 59 rows - 1 Africa country/area - 2021 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 59 rows from Cote d'Ivoire DataFair, covering Sites De Vaccination De Covid 19 Dans Le District D Abidja. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-cote-d-ivoire-sites-de-vaccination-de-covid-19-dans-le-district-d-abidja-5bbc9c04.imagetabular-classificationn<1K0 likes41 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.