datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
minimal-en-corpus-5b
Minimal EN Corpus 5B
An English-language pretraining corpus prepared for controlled experiments with approximately 125M-parameter GPT-2 models based on karpathy/nanoGPT.
The name refers to the approximately 5B-token mixture selected before final BPE tokenization. With the included 12,288-token BPE tokenizer, the packaged nanoGPT training split contains 5,396,605,407 tokens.
Contents
The dataset provides both reusable source text and ready-to-train nanoGPT… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b.pa-warm-start-sft-medium-5b-mix
geodesic-research/pa-warm-start-sft-medium-5b-mix
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/pa-warm-start-sft-medium-5b-mix", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-medium-5b-mix.icrm-hitek-full-db-mixed-5b
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/idkdd/icrm-hitek-full-db-mixed-5b.fineweb-5Bfineweb-edu-dedup-5Bd24-midtrain-olmo3-5b
d24 Midtrain — OLMo-3 Dolmino (5B, chunked)
A 5.00B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix, built by
taking a fraction of each component of
allenai/dolma3_dolmino_mix-100B-1025.
Unlike the smaller d24-midtrain-olmo3
(which dropped documents over 2048 tokens — silently zeroing the long reasoning-trace and PDF
components), this build chunks long documents into 2048-token windows (decoded back to text),
so every component keeps ~100% of its tokens and reaches… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b.InfLLM-V2-data-5B
InfLLM-V2 Long-Context Training Dataset with 5B Tokens
Project Links: [Paper] [InfLLM-V2 Models] [CUDA Kernel Code]
🚀 About InfLLM-V2
InfLLM-V2 is a native sparse attention framework designed for the efficient processing of long-sequence texts. Its core advantage is the ability to maintain high performance comparable to dense attention in short-text scenarios—without any extra parameters—while seamlessly switching to a sparse mode for long-text scenarios, achieving… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/InfLLM-V2-data-5B.d24-midtrain-olmo3-5b-wholedoc
d24 Midtrain — OLMo-3 Dolmino (5B, whole-doc)
A 4.71B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix at its exact
component proportions, built by taking a fraction of each component of
allenai/dolma3_dolmino_mix-100B-1025.
Documents are kept whole — no length filter, no chunking.
Why whole-doc: for packed pretrain/midtrain the trainer (e.g. Megatron preprocess_data --append-eod
GPTDataset) already concatenates documents and slices them into context-length windows… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b-wholedoc.africa-sudan-sudan-environment-5b08204c
Sudan - Environment | Africa (Sudan official open data)
4,043 rows - 1 Africa country - 1961-2025 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Sudan as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Sudan - Environment
Publisher: World Bank Group
Resource: Environment… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-sudan-sudan-environment-5b08204c.med-domain-5brmfs-wm-v3-jenkins5b
RMFS World Model v3 — jenkins5-b1
Egocentric driving data for a single robot (the ego) in a simulated Robotic
Mobile Fulfillment System (RMFS) warehouse, collected from
RAWSim-O via a custom Gym server.
Intended for training World Models-style
V (VAE) + M (MDN-RNN) + C (controller) stacks where the controller is trained
entirely inside the learned dream.
At a glance
Episodes
120
Frames
512,734
Hops (macro-steps)
28,800
Cameras
4 x 96x96 RGB… See the full description on the dataset page: https://huggingface.co/datasets/luckysantoso/rmfs-wm-v3-jenkins5b.starcoder-python5b5b gpt2 tokens
NEW_qwen2_5_MATH_1_5b_grpo_reg_grpo_bce_4dclm_seed_5b_tanishqInfLLM-V2-data-5B-v2
InfLLM-V2 Long-Context Training Dataset with 5B Tokens
Project Links: [Paper] [InfLLM-V2 Models] [CUDA Kernel Code]
🚀 About InfLLM-V2
InfLLM-V2 is a native sparse attention framework designed for the efficient processing of long-sequence texts. Its core advantage is the ability to maintain high performance comparable to dense attention in short-text scenarios—without any extra parameters—while seamlessly switching to a sparse mode for long-text scenarios, achieving… See the full description on the dataset page: https://huggingface.co/datasets/elonmuskceo/InfLLM-V2-data-5B-v2.NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_4NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_2NEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gpg_bce_5NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_5r8-eval-suite-5bucket
⚠️ CRITICAL: Ollama Inference Flag Required for derived models
If you train or serve any Qwen3.5-9B-derived model from this lineage via Ollama,
you MUST pass "think": false in /api/chat requests for chat / instruction following / tool use.
The qwen3.5 RENDERER auto-injects <think> tags causing 25-46% empty-answer rates without this flag.
See dataset cudabenchmarktest/r9-research-framework/_OLLAMA_INFERENCE_WARNING.md for the full lesson learned.
R8/R9 Five-Bucket… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r8-eval-suite-5bucket.InfLLM-V2-data-5B
InfLLM-V2 Long-Context Training Dataset with 5B Tokens
Project Links: [Paper] [InfLLM-V2 Models] [CUDA Kernel Code]
🚀 About InfLLM-V2
InfLLM-V2 is a native sparse attention framework designed for the efficient processing of long-sequence texts. Its core advantage is the ability to maintain high performance comparable to dense attention in short-text scenarios—without any extra parameters—while seamlessly switching to a sparse mode for long-text scenarios, achieving… See the full description on the dataset page: https://huggingface.co/datasets/elonmuskceo/InfLLM-V2-data-5B.GPT2-Large-Llama3-8B-fineweb-256-5Btokens
Source:
Modified from HuggingFaceFW/fineweb Sample-10BT subset.
The preprocessed script files are in the data directory. You can download it and run:
python fineweb.py --segment_length 512 --nproc 32 --batch_size 1024 --save_path ./fineweb10B/save/
Dataset Owner(s):
Individual: BroAlanTaps
License/Terms of Use:
odc-by: Open Data Commons License Attribution family
Intended Usage:
This dataset is intended to be used with research or… See the full description on the dataset page: https://huggingface.co/datasets/BroAlanTaps/GPT2-Large-Llama3-8B-fineweb-256-5Btokens.bright-passage-index-gte_qwen2-1_5bNEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gspo_bce_4atlas9_5beh_sequential_sdfNEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gspo_bce_3codelion_all-5BNEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gspo_bce_5NEW_qwen2_5_MATH_1_5b_grpo_reg_grpo_bce_2africa-cote-d-ivoire-sites-de-vaccination-de-covid-19-dans-le-district-d-abidja-5bbc9c04
Sites De Vaccination De Covid 19 Dans Le District D Abidja | Africa (Cote d'Ivoire DataFair)
59 rows - 1 Africa country/area - 2021 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 59 rows from Cote d'Ivoire DataFair, covering Sites De Vaccination De Covid 19 Dans Le District D Abidja. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-cote-d-ivoire-sites-de-vaccination-de-covid-19-dans-le-district-d-abidja-5bbc9c04.
