datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
d24-midtrain-olmo3-5b
d24 Midtrain — OLMo-3 Dolmino (5B, chunked)
A 5.00B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix, built by
taking a fraction of each component of
allenai/dolma3_dolmino_mix-100B-1025.
Unlike the smaller d24-midtrain-olmo3
(which dropped documents over 2048 tokens — silently zeroing the long reasoning-trace and PDF
components), this build chunks long documents into 2048-token windows (decoded back to text),
so every component keeps ~100% of its tokens and reaches… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b.d24-midtrain-olmo3-5b-wholedoc
d24 Midtrain — OLMo-3 Dolmino (5B, whole-doc)
A 4.71B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix at its exact
component proportions, built by taking a fraction of each component of
allenai/dolma3_dolmino_mix-100B-1025.
Documents are kept whole — no length filter, no chunking.
Why whole-doc: for packed pretrain/midtrain the trainer (e.g. Megatron preprocess_data --append-eod
GPTDataset) already concatenates documents and slices them into context-length windows… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b-wholedoc.starcoder-python5b5b gpt2 tokens
r8-eval-suite-5bucket
⚠️ CRITICAL: Ollama Inference Flag Required for derived models
If you train or serve any Qwen3.5-9B-derived model from this lineage via Ollama,
you MUST pass "think": false in /api/chat requests for chat / instruction following / tool use.
The qwen3.5 RENDERER auto-injects <think> tags causing 25-46% empty-answer rates without this flag.
See dataset cudabenchmarktest/r9-research-framework/_OLLAMA_INFERENCE_WARNING.md for the full lesson learned.
R8/R9 Five-Bucket… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r8-eval-suite-5bucket.qwen2-5-1-5b-blindspots
Qwen2.5-1.5B Blind Spots Dataset
Overview
This dataset contains failure cases collected while testing the open model Qwen/Qwen2.5-1.5B. I selected this model because it is an open, general-purpose base model within the required parameter range and easy to load on a free GPU notebook environment.
The goal of this dataset is to document cases where the model makes incorrect predictions or shows important blind spots. I focused on diverse failure types rather than repeating… See the full description on the dataset page: https://huggingface.co/datasets/Akil139/qwen2-5-1-5b-blindspots.canada-dataset-5b
🇨🇦 Canada Web Text — 5B-token Sample 🍁
A 5-billion-token sample of cleaned Canada-attributed web text derived from Common Crawl. This sample was produced as part of the JoeyLLM project's ongoing research into regional English language datasets. 🌐
This dataset is intended as a large-scale Canadian English web-text corpus for language-model pre-training, continued pre-training, data inspection, and regional English research.
📊 Dataset Summary 📌
Property… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/canada-dataset-5b.my-distiset-d2a39c5b
Dataset Card for my-distiset-d2a39c5b
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/phydmod/my-distiset-d2a39c5b/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/phydmod/my-distiset-d2a39c5b.new-zealand-dataset-5b
🇳🇿 New Zealand Web Text — 5B-token Sample 🌿
A 5-billion-token representative sample of a much larger cleaned New Zealand web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐
The full 60.7B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/new-zealand-dataset-5b.australian-dataset-5b
🇦🇺 Australian Web Text — 5B-token Sample 🦘
A 5-billion-token Australian web-text dataset created for the JoeyLLM project. This dataset was sampled from the filtered Australian corpus produced by the JoeyLLM sovereign corpus pipeline. 🌐
The purpose of this dataset is to provide a large-scale Australian text corpus for GPT-style language-model pre-training, continued pre-training, data inspection, and research into regional English language models.
📊 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/australian-dataset-5b.uk-dataset-5b
🇬🇧 UK Web Text — 5B-token Sample
A 5-billion-token sample of cleaned United Kingdom web text derived from Common Crawl. This sample was produced as part of the JoeyLLM project's ongoing research into regional English language datasets. 🌐
This dataset is intended as a large-scale UK-attributed English web-text corpus for language-model pre-training, continued pre-training, data inspection, and regional English research.
📊 Dataset Summary 📌
Property
This… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/uk-dataset-5b.
