CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thu-pacman /Puro-2B Puro-2B Pretraining Data: The Recipe Behind a 2B Model This is the materialized pretraining data release for Puro-2B-Base, a 2B base model trained from scratch on consumer-grade RTX 5090 GPUs. The repository contains the component-level data pools used to construct the two Puro-2B pretraining phases, together with the tokenizer used for token accounting. It is organized for inspection, selective streaming, and recipe reconstruction rather than as a small train/test… See the full description on the dataset page: https://huggingface.co/datasets/thu-pacman/Puro-2B.texttext-generation100M<n<1B8 likes7.3k downloads1d agoHugging Face02thu-pacman /PCMind-2.1-Kaiyuan-2B This repository contains the complete pretraining dataset for PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model. Overview The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains: English: General English text Chinese: General Chinese text Code: Programming and code-related content Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/thu-pacman/PCMind-2.1-Kaiyuan-2B.texttext-generation1B<n<10B5 likes2.4k downloads10mo agoHugging Face03Lxd99 /PCMind-2.1-Kaiyuan-2B-phase1-part1-2 This repository contains the complete pretraining dataset for PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model. Overview The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains: English: General English text Chinese: General Chinese text Code: Programming and code-related content Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-2.texttext-generation100M<n<1B0 likes449 downloads6mo agoHugging Face04Lxd99 /PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323 This repository contains the complete pretraining dataset for PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model. Overview The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains: English: General English text Chinese: General Chinese text Code: Programming and code-related content Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323.texttext-generation100M<n<1B0 likes442 downloads6mo agoHugging Face05JWei05 /gemma4-e2b-base-topk128-hf-overlay-v128-seed42 Gemma 4 E2B base top-k-128 HF training overlay This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from JWei05/gemma4-e2b-base-topk128-traces, but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training engine. This repository is a reproducibility artifact for the corresponding distillation run. It is not a new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.tabulartext-generation10K<n<100K0 likes312 downloads2mo agoHugging Face06headwaterai /headwater-b2b-data-purification-ai-ingestion-implementation-kit HeadWater B2B Data Purification & AI Ingestion Implementation Kit He will not suffer thy foot to be moved: he that keepeth thee will not slumber. Psalm 121:3 YOU CAN HAVE IT NOW. About This Kit The Headwater B2B Data Purification & AI Ingestion Implementation Kit is engineered for one primary purpose: to eliminate development delays and buy back your operational momentum. Instead of wasting weeks building foundational data plumbing from scratch… See the full description on the dataset page: https://huggingface.co/datasets/headwaterai/headwater-b2b-data-purification-ai-ingestion-implementation-kit.text-generation0 likes305 downloads12h agoHugging Face07SlayerLab /gollem-corpus-2b-pl GoLLeM Corpus 2B PL Dokładny korpus treningowy polskiego modelu bazowego SlayerLab/GoLLeM-110M-PL-v3 (oraz v2) — ten sam zbiór, po którym model przeszedł dwie epoki. Publikujemy go, aby każdy mógł odtworzyć trening od zera na własnym tokenizerze. Jak powstał ten plik. Korpus odzyskano przez zdekodowanie stokenizowanego checkpointu treningowego (byte-level BPE dynaword-32k, round-trip bezstratny; granice dokumentów = token <|endoftext|>). To jest dokładnie tekst, który model… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-2b-pl.tabulartext-generation1M<n<10M1 likes287 downloads25d agoHugging Face08F555 /qwen3.5-2b-base-blind-spots Qwen3.5-2B-Base — Blind Spot Analysis (Text + Vision) Model Tested Field Value Model Qwen/Qwen3.5-2B-Base Parameters 2.27 B (2,274 M per HF metadata) Architecture Hybrid Gated-DeltaNet (dense FFN) — 24 LM layers (18 DeltaNet + 6 full-attention), ViT vision encoder Type Pre-trained base model (not instruction-tuned) Context 262 144 tokens Modalities Text + Vision (early-fusion multimodal) Key Contributions Only multimodal… See the full description on the dataset page: https://huggingface.co/datasets/F555/qwen3.5-2b-base-blind-spots.imagetext-generationn<1K0 likes141 downloads6mo agoHugging Face09exnivo /tinybrain-pretrain-corpus-2b TinyBrain Pretrain Corpus 2B A mixed-source English pretraining corpus for training small language models. TinyBrain Pretrain Corpus 2B is a mixed-source dataset built for pretraining small causal language models, especially the TinyBrain-100M Base model. The dataset combines educational text, factual/wiki-style text, math reasoning data, Python code-summary data, clean web text, and conversation-style data. It is designed to give small models a useful general foundation… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-pretrain-corpus-2b.texttext-generation1M<n<10M1 likes126 downloads3mo agoHugging Face10Ba2han /ultra-fineweb-tokenized-2B Ultra-FineWeb Tokenized 2B Status: complete A tokenized subset streamed from [openbmb/Ultra-FineWeb] using its en split. The Parquet data has exactly one column: input_ids (list<int32>). Processing Source revision: 7ddd4170ce03e0afbd7d9b80d4bc0b8eebf877e4 Tokenizer: Ba2han/TR_CPT1 Tokenizer revision: d32fb40763740cca0a7c04b4549d2129edeafa9a Source field: content Filter: score > 0.75 Stored sequence length: 50 to 3,000 tokens, inclusive Target: approximately 2,000… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/ultra-fineweb-tokenized-2B.text-generation1M<n<10M0 likes115 downloads2mo agoHugging Face11enaix /ml2b ML2B: Multi-Lingual ML Benchmark For AutoML This repository provides the dataset for ML2B (Multi-Lingual ML Benchmark for AutoML), the first benchmark for evaluating multilingual machine learning (ML) code generation. Presented in the paper ML2B: Multi-Lingual ML Benchmark For AutoML, ML2B consists of 30 Kaggle competitions translated into 13 natural languages. It covers tabular, text, and image data types, and includes structured metadata and validated human-reviewed… See the full description on the dataset page: https://huggingface.co/datasets/enaix/ml2b.imagetext-generation0 likes75 downloads23d agoHugging Face12arjhinety /OpenGrad-Qwen3.5-2B-M0-SFT-CorpusV1-evaluation OpenGrad — Qwen3.5-2B, M0 SFT on corpus v1: evaluation record This repository holds the evaluation evidence for one OpenGrad experiment, qwen35_2b_m0_sft_full_v3: a full-parameter supervised fine-tuning run of Qwen/Qwen3.5-2B on the published OpenGrad ToolPolicy Canonical v1 corpus. There are no model weights here, and none exist. Every checkpoint this run produced was deleted from local storage before it was uploaded, and none of them can be recovered. This repository is what… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-Qwen3.5-2B-M0-SFT-CorpusV1-evaluation.text-generation0 likes71 downloads1d agoHugging Face13wisdompan /qwen35-2b-personal-training-data Qwen3.5-2B Three-Domain Training Data A reproducible training-data release assembled and processed by wisdompan for Qwen3.5-2B experiments across mathematics, code, and instruction following. Dataset configurations Configuration Purpose Train rows Validation rows full_mix Unified three-domain student training 86,931 3 teacher_math Mathematics teacher training 17,917 1 teacher_code Code teacher training 23,667 1 teacher_if Instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/wisdompan/qwen35-2b-personal-training-data.texttext-generation100K<n<1M0 likes67 downloads17d agoHugging Face14Gingiris /gingiris-b2b-growth 📈 Gingiris B2B SaaS Growth Playbook Scale your B2B SaaS from PMF to $10M ARR — PLG/SLG motion selection, affiliate marketing, channel partnerships. Battle-tested with HeyGen, Deel, Vercel, Supabase, and AWS case studies. Built by Iris (生姜iris), Forbes Asia 30 Under 30. English | 中文 | 日本語 | 한국어 📦 Install npx skills add Gingiris-1031/gingiris-b2b-growth Then ask your AI agent: "We're a B2B SaaS at $5k MRR — how do I scale to $50k?" or "Should we switch from… See the full description on the dataset page: https://huggingface.co/datasets/Gingiris/gingiris-b2b-growth.text-generationn<1K0 likes58 downloads20d agoHugging Face15open-athena /Snowball-67B-A2B-RLVR1-Repro-Data Snowball 67B-A2B RLVR1 data These are the exact Parquet inputs retained for the Snowball 67B-A2B sync and async RLVR1 experiments on Iris cw-rno2a in September 2026. The data was selected from the skyrl_gym route of a TaskTrove conversion of the public NVIDIA Nemotron RL Ultra training blend, preserving source order and holding out the last 100 selected rows. See provenance.json for the local conversion and filtering record. The original TaskTrove release is also public.… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-RLVR1-Repro-Data.texttext-generation10K<n<100K0 likes58 downloads7d agoHugging Face16Jurgen1161 /synthetic-b2b-saas-support-dialogues-sample Synthetic B2B SaaS Support Dialogues (Sample) Free sample: 100 dialogues from a larger dataset of 484 synthetic customer support conversations for B2B SaaS products. What's inside 100 complete dialogues (6–8 messages each) 7 issue categories: auth, billing, integration, data, account, technical, onboarding Rich metadata: resolution_status, customer_sentiment, agent_actions, escalation_needed Realistic technical details: error codes, URLs, button names, account… See the full description on the dataset page: https://huggingface.co/datasets/Jurgen1161/synthetic-b2b-saas-support-dialogues-sample.texttext-generationn<1K0 likes50 downloads8d agoHugging Face17dmnsh /caliber-extension-gemma4-e2b-grpo-rollouts CALIBER Extension — Gemma4-E2B GRPO Rollouts Training rollouts from matched GRPO arms on google/gemma-4-E2B-it (new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps). Subsets subset arm τ prior rows mean reward_total accuracy full schema caliber vanilla CALIBER 0.0 — 1600 2.298 0.514 0.664 mink Min-K% prior 1.0 mink_0.2 4800 2.506 0.520 0.680 minkpp Min-K++% prior 1.0 minkpp_0.2 4800 2.637 0.541 0.726 Load: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.tabulartext-generation10K<n<100K0 likes46 downloads13d agoHugging Face18asingh15 /qwen35-2b-tool-use-qwen36-27b-curation-candidates Full candidate collections: 2B tool use + 27B data curation This public Dataset contains two complete, unredacted, exact-40 candidate collections: Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and 233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL, Spider, and TravelPlanner. Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021 targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.tabulartext-generation100K<n<1M0 likes45 downloads1mo agoHugging Face19Solshine /nla-gemma4e2b-relabel-v1-eval Gemma-4-E2B layer-23 evaluation set, relabeled (v1), with contamination flags The 580-document evaluation pool on which every activation-verbalizer result in this project is scored, with each row's evaluation text rewritten from a topic summary to a feature-attribution label. The activations are byte-identical to the original evaluation set; only the text column changed, and the original text is preserved. This pool is not disjoint from the training corpus. Read this… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-relabel-v1-eval.texttext-generationn<1K0 likes44 downloads7d agoHugging Face20kiddothe2b /synthetic_polistance Fully Synthetic Prompts for LLM Political Stance Detection All resources developed in the article "Templated or fully Synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance" (Chalkidis, 2026). Paper Abstract Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions—originally designed for humans, and thus lacks the realism and nuance of human-AI… See the full description on the dataset page: https://huggingface.co/datasets/kiddothe2b/synthetic_polistance.tabulartext-generation1K<n<10K0 likes43 downloads1mo agoHugging Face21liswei /Taiwan-Text-Excellence-2Bgated High quality corpus for Taiwanese culture and Traditional Chinese Taiwan Text Excellence (TTE) Contains high quality news and articles in Traditional Chinese. The data processing pipeline is optimized for LLM performance. Is de-duplicated and cleaned using both rule-based and learning-based filters. E.g., urls/emails/html tags/abnormal characters are cleaned, and numbers (full-width or half-width) are normalized. Contains ~2 billion tokens, measured using BPE tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/liswei/Taiwan-Text-Excellence-2B.texttext-generation1M<n<10M22 likes40 downloads2y agoHugging Face22esherialabs /saferide-gemma-4-e2b-v058-original-419806-training-data SafeRide Synthetic Bilingual Safety Guidance Dataset v0.5.8 This research and development dataset contains synthetic English and Kiswahili chat conversations. It was designed to help a language model practice cautious, agency-preserving safety guidance, useful refusal behavior, and responses that avoid inventing facts. It contains no real survivor reports or production records. The frozen dataset is publicly available under Creative Commons Attribution 4.0 International (CC BY… See the full description on the dataset page: https://huggingface.co/datasets/esherialabs/saferide-gemma-4-e2b-v058-original-419806-training-data.texttext-generation1K<n<10K0 likes39 downloads1mo agoHugging Face23Solshine /nla-gemma4e2b-relabel-v1-corpus Gemma-4-E2B layer-23 activation corpus, relabeled (v1) 1356 training rows for an activation verbalizer. Each row pairs a residual-stream activation captured at layer 23 of google/gemma-4-E2B with a natural-language label describing what the model must have integrated at that position to predict its next token. This is the training set behind Solshine/gemma-4-e2b-nla-L23-av-priordev-relabel-v1-wd3. Why it exists An audit of the previous version of this corpus found… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-relabel-v1-corpus.tabulartext-generation1K<n<10K0 likes38 downloads7d agoHugging Face24asingh15 /qwen35-2b-tool-use-candidates Qwen3.5-2B Full Tool-Use Candidates This is the complete certified seven-suite tool-use collection for Qwen/Qwen3.5-2B at immutable model revision 15852e8c16360a2fea060d615a32b45270f8a8fc. 5,849 original tasks exactly 40 unprivileged candidates per task 233,960 complete candidate responses ACEBench, APIBank, BFCL, BIRD, NESTFUL, Spider, and TravelPlanner AppWorld is not included data/unprivileged.jsonl is a byte-for-byte copy of the certified collection. Original task IDs… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-candidates.texttext-generation1K<n<10K0 likes35 downloads1mo agoHugging Face25daipham31 /qwen3.5-2B-vi-query Vietnamese Medical Query Normalization / Expansion / Routing Pack (v3) 1162 synthetic ChatML examples for fine-tuning a small Vietnamese model (target: Qwen/Qwen3.5-2B, trained with Unsloth) to turn a raw, everyday Vietnamese medical query into structured JSON: normalized query, intent, entities, must-preserve tokens, lexical/semantic query variants, and a retrieval-routing hint, for a downstream medical RAG system. The model does not answer medical questions. It only normalizes… See the full description on the dataset page: https://huggingface.co/datasets/daipham31/qwen3.5-2B-vi-query.texttext-generation1K<n<10K0 likes35 downloads3d agoHugging Face26mags0ft /Gemma-4-E2B-SSFT Gemma-4-E2B-SSFT This is a dataset that has been created using SSFT (simple-sft), a synthetic data generation tool written by me. It contains a few hundred samples for testing. Try it out yourself! texttext-generationn<1K1 likes33 downloads2mo agoHugging Face27zcamz /ai-vs-human-google-gemma-2-2b-it AI vs Human dataset on the CNN Daily mails Dataset Description This dataset showcases pairs of truncated articles and their respective completions, crafted either by humans or an AI language model. Each article was randomly truncated between 25% and 50% of its length. The language model was then tasked with generating a completion that mirrored the characters count of the original human-written continuation. Data Fields 'human': The original human-authored… See the full description on the dataset page: https://huggingface.co/datasets/zcamz/ai-vs-human-google-gemma-2-2b-it.texttext-classification1K<n<10K1 likes30 downloads2y agoHugging Face28FreeAIn /Pwen3.5_2B_Python_Finetune Pwen3.5-2B-Coding-Finetune Pwen 3.5 2B Coding Dataset A high-quality instruction dataset for fine-tuningQwen3.5-2B into a concise coding assistant Created by Pavel Hanzel Overview Pwen3.5-2B-Coding-Finetune is an instruction tuning dataset designed to transform Qwen3.5-2B into a practical programming assistant. The dataset focuses on: Python programming Debugging Code explanations Development workflows AI/LLM usage Direct technical… See the full description on the dataset page: https://huggingface.co/datasets/FreeAIn/Pwen3.5_2B_Python_Finetune.texttext-generationn<1K0 likes30 downloads3mo agoHugging Face29science-of-finetuning /ultrachat_200k_generated_gemma-2-2b-itThis dataset contains 512 answers generated by the gemma-2-2b-it model on a subset of the ultrachat 200k test_sft dataset using greedy decoding. The subset was generated by filtering out conversations that were >= 1024 - 128 tokens long, and answers were cut off at each batch after 1024 - min(batch_prompt_lengths) generated tokens, such that each answer is at most 128 tokens long. The generated answers are 200k tokens so 390 tokens (~300 words or 2/3 pages) on average. texttext-generationn<1K0 likes24 downloads2y agoHugging Face30k-imtz /youtu-llm-2b-base-blind-spots Youtu-LLM-2B-Base Blind Spots Evaluation Dataset This dataset contains 75 evaluation prompts used to analyze the failure modes of tencent/Youtu-LLM-2B-Base, a 1.96B parameter dense base language model released on December 31, 2025. Each row includes the input prompt, the expected answer, and the model’s generated output obtained during inference on a Google Colab T4 GPU. The prompts span 13 broad categories including arithmetic, logic, multilingual generation, instruction following… See the full description on the dataset page: https://huggingface.co/datasets/k-imtz/youtu-llm-2b-base-blind-spots.texttext-generationn<1K0 likes24 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.