CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sfanm /d24-midtrain-olmo3-5b d24 Midtrain — OLMo-3 Dolmino (5B, chunked) A 5.00B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix, built by taking a fraction of each component of allenai/dolma3_dolmino_mix-100B-1025. Unlike the smaller d24-midtrain-olmo3 (which dropped documents over 2048 tokens — silently zeroing the long reasoning-trace and PDF components), this build chunks long documents into 2048-token windows (decoded back to text), so every component keeps ~100% of its tokens and reaches… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b.texttext-generation10M<n<100M0 likes346 downloads3mo agoHugging Face02sfanm /d24-midtrain-olmo3-5b-wholedoc d24 Midtrain — OLMo-3 Dolmino (5B, whole-doc) A 4.71B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix at its exact component proportions, built by taking a fraction of each component of allenai/dolma3_dolmino_mix-100B-1025. Documents are kept whole — no length filter, no chunking. Why whole-doc: for packed pretrain/midtrain the trainer (e.g. Megatron preprocess_data --append-eod GPTDataset) already concatenates documents and slices them into context-length windows… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b-wholedoc.texttext-generation10M<n<100M0 likes292 downloads3mo agoHugging Face03lparkourer10 /starcoder-python5b5b gpt2 tokens tabulartext-generation1M<n<10M0 likes205 downloads2y agoHugging Face04cudabenchmarktest /r8-eval-suite-5bucket ⚠️ CRITICAL: Ollama Inference Flag Required for derived models If you train or serve any Qwen3.5-9B-derived model from this lineage via Ollama, you MUST pass "think": false in /api/chat requests for chat / instruction following / tool use. The qwen3.5 RENDERER auto-injects <think> tags causing 25-46% empty-answer rates without this flag. See dataset cudabenchmarktest/r9-research-framework/_OLLAMA_INFERENCE_WARNING.md for the full lesson learned. R8/R9 Five-Bucket… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r8-eval-suite-5bucket.tabulartext-generationn<1K0 likes131 downloads5mo agoHugging Face05Akil139 /qwen2-5-1-5b-blindspots Qwen2.5-1.5B Blind Spots Dataset Overview This dataset contains failure cases collected while testing the open model Qwen/Qwen2.5-1.5B. I selected this model because it is an open, general-purpose base model within the required parameter range and easy to load on a free GPU notebook environment. The goal of this dataset is to document cases where the model makes incorrect predictions or shows important blind spots. I focused on diverse failure types rather than repeating… See the full description on the dataset page: https://huggingface.co/datasets/Akil139/qwen2-5-1-5b-blindspots.texttext-generationn<1K0 likes10 downloads7mo agoHugging Face06JoeyLLM /canada-dataset-5bgated 🇨🇦 Canada Web Text — 5B-token Sample 🍁 A 5-billion-token sample of cleaned Canada-attributed web text derived from Common Crawl. This sample was produced as part of the JoeyLLM project's ongoing research into regional English language datasets. 🌐 This dataset is intended as a large-scale Canadian English web-text corpus for language-model pre-training, continued pre-training, data inspection, and regional English research. 📊 Dataset Summary 📌 Property… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/canada-dataset-5b.texttext-generation1M<n<10M0 likes9 downloads4mo agoHugging Face07phydmod /my-distiset-d2a39c5b Dataset Card for my-distiset-d2a39c5b This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/phydmod/my-distiset-d2a39c5b/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/phydmod/my-distiset-d2a39c5b.texttext-generationn<1K0 likes7 downloads2y agoHugging Face08JoeyLLM /new-zealand-dataset-5bgated 🇳🇿 New Zealand Web Text — 5B-token Sample 🌿 A 5-billion-token representative sample of a much larger cleaned New Zealand web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐 The full 60.7B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/new-zealand-dataset-5b.tabulartext-generation1M<n<10M0 likes5 downloads4mo agoHugging Face09JoeyLLM /australian-dataset-5bgated 🇦🇺 Australian Web Text — 5B-token Sample 🦘 A 5-billion-token Australian web-text dataset created for the JoeyLLM project. This dataset was sampled from the filtered Australian corpus produced by the JoeyLLM sovereign corpus pipeline. 🌐 The purpose of this dataset is to provide a large-scale Australian text corpus for GPT-style language-model pre-training, continued pre-training, data inspection, and research into regional English language models. 📊 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/australian-dataset-5b.tabulartext-generation1M<n<10M0 likes4 downloads4mo agoHugging Face10JoeyLLM /uk-dataset-5bgated 🇬🇧 UK Web Text — 5B-token Sample A 5-billion-token sample of cleaned United Kingdom web text derived from Common Crawl. This sample was produced as part of the JoeyLLM project's ongoing research into regional English language datasets. 🌐 This dataset is intended as a large-scale UK-attributed English web-text corpus for language-model pre-training, continued pre-training, data inspection, and regional English research. 📊 Dataset Summary 📌 Property This… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/uk-dataset-5b.texttext-generation1M<n<10M0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.