CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MaxDevv /Qwen3.8-27B-Distill-1M-3.12B-Tokens Qwen3.8-27B-Distill-1M-4.83B-Tokens A unified, globally deduplicated, large-scale supervised distillation corpus built from 992,318 conversations generated by Qwen/Qwen3.8-27B, containing 4,834,771,862 target output tokens (3,570,459,498 reasoning tokens + 1,264,312,364 final response tokens) and 5,104,980,053 total sequence tokens. 1. Dataset Overview This dataset merges, aligns, and deduplicates the two primary high-quality Qwen3.8-27B generation corpora on… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/Qwen3.8-27B-Distill-1M-3.12B-Tokens.texttext-generation100K<n<1M1 likes316 downloads27d agoHugging Face02procmarco /fpabl1-arm-b-fp-tokens-48k fpabl1-arm-b-fp-tokens-48k Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-b-fp. Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121). Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total. Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks Source token composition: fp_en: 1,000,000,000 fp_ita:… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-b-fp-tokens-48k.tabularn<1K0 likes178 downloads3mo agoHugging Face03Pacific-i64 /data-32k-200b-tokens TR-HASH 32K · 200B Token Mixture Pretokenized training mixture for compact TR-HASH language-model research. The Dataset Viewer displays one summary row per source. The actual training data is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin. Each sequence contains 1,024 token IDs produced by the project 32K tokenizer. Mixture Source Weight Training tokens DCLM 45% 90B FineWeb-Edu deduplicated 30% 60B Stack-Edu 10% 20B… See the full description on the dataset page: https://huggingface.co/datasets/Pacific-i64/data-32k-200b-tokens.tabulartext-generationn<1K1 likes114 downloads1mo agoHugging Face04ZhiyuanQiu /RAW20seul_2316273_tokenstext100K<n<1M0 likes67 downloads4y agoHugging Face05ZhiyuanQiu /RAW10seul_1262172_tokenstext100K<n<1M0 likes65 downloads4y agoHugging Face06xiaobo6668 /math-soft-tokens Math Soft Tokens Dataset Contains training steps: numinamath15_step_11_fixed. text10K<n<100K0 likes59 downloads9mo agoHugging Face07Data-Elite /French-Expert-SFT-81M-Tokens French Expert SFT Corpus (81M Tokens) 🎯 Description Ce dataset est un corpus de haute qualité conçu pour le Supervised Fine-Tuning (SFT). Il a été constitué par un moteur de recherche thématique profond (deep-crawl) ciblant les domaines de haute expertise technique et juridique française. 📊 Statistiques Clés Nombre total de pépites (Samples) : 456,863 Volume estimé : ~81 Millions de Tokens Taille moyenne par entrée : 629 caractères Qualité : 0% doublons… See the full description on the dataset page: https://huggingface.co/datasets/Data-Elite/French-Expert-SFT-81M-Tokens.text10K<n<100K0 likes50 downloads9mo agoHugging Face08TokenBender /unnatural_code_instructions_20M_tokens_separatetext100K<n<1M4 likes41 downloads3y agoHugging Face09aweffr /chinese-novel-continuation-precise-tokens 中文小说续写精确 Token 长度数据集 基于金庸《神雕侠侣》的高精度中文文本生成数据集,专为 GPU 内存测试和序列长度性能分析而设计。 数据集特点 高精度:99.4%+ 的目标 token 长度准确率 多种长度:5 个变体(1024、2048、4096、8192、16384 tokens) 统一格式:Alpaca 格式的小说续写任务 质量控制:95%+ 样本在目标长度的 ±2% 范围内 数据集统计 目标 Tokens 实际平均 准确率 样本数 文件大小 ±1% 内样本 ±2% 内样本 1024 1017.9 99.4% 800 2.5MB 739 775 2048 2037.7 99.5% 800 4.9MB 764 795 4096 4078.1 99.6% 800 9.8MB 792 800 8192 8158.1 99.6% 800 19.5MB 800 800 16384 16317.6 99.6% 800 39.1MB 800 800… See the full description on the dataset page: https://huggingface.co/datasets/aweffr/chinese-novel-continuation-precise-tokens.text1K<n<10K4 likes41 downloads1y agoHugging Face10AETHORIA-AI /data-32k-200b-tokens TR-HASH 32K · 200B Token Mixture Pretokenized training mixture for compact TR-HASH language-model research. The Dataset Viewer displays one summary row per source. The actual training data is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin. Each sequence contains 1,024 token IDs produced by the project 32K tokenizer. Mixture Source Weight Training tokens DCLM 45% 90B FineWeb-Edu deduplicated 30% 60B Stack-Edu 10% 20B… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/data-32k-200b-tokens.tabulartext-generationn<1K0 likes40 downloads1mo agoHugging Face11pedrodev2026 /pedro-open-dataset-max-512-tokenstexttext-generation10K<n<100K0 likes35 downloads7mo agoHugging Face12pedrodev2026 /pedro-open-dataset-max-512-tokens-25ktexttext-generation10K<n<100K0 likes35 downloads7mo agoHugging Face13ZhiyuanQiu /RAW15seul_1834048_tokenstext100K<n<1M0 likes33 downloads4y agoHugging Face14cotarc0 /model_20_tokens_10_specstabularn<1K0 likes33 downloads1y agoHugging Face15cotarc0 /model_20_tokens_3_specstabularn<1K0 likes28 downloads1y agoHugging Face16Lin1557 /Critical-Tokens-Matter-Train-Datatext10K<n<100K0 likes27 downloads1y agoHugging Face17cotarc0 /model_20_tokens_50_specstabularn<1K0 likes24 downloads1y agoHugging Face18pedrodev2026 /pedro-open-dataset-max-512-tokens-10ktexttext-generation10K<n<100K0 likes23 downloads7mo agoHugging Face19Ma7ee7 /Instruct-Data-7M-Tokenstext1K<n<10K0 likes22 downloads2mo agoHugging Face20cotarc0 /model_20_tokens_100_specstabularn<1K0 likes21 downloads1y agoHugging Face21cotarc0 /model_20_tokens_20_specstabularn<1K0 likes20 downloads1y agoHugging Face22cotarc0 /model_20_tokens_200_specstabularn<1K0 likes17 downloads1y agoHugging Face23cotarc0 /model_20_tokens_5_specstabularn<1K0 likes17 downloads1y agoHugging Face24pedrodev2026 /opencodegeneticinstruct-max-512-tokens-10ktext10K<n<100K0 likes14 downloads6mo agoHugging Face25fxmeng /transmla_pretrain_100m_tokenstext10K<n<100K0 likes12 downloads1y agoHugging Face26nassimjp /pashto-warmup-tokens Pashto Warmup Tokens Dataset This dataset contains a curated, deduplicated collection of high-quality, contextually accurate Pashto linguistic examples. It maps structural language tasks directly to the most critical vocabulary tokens in Pashto, providing a reliable corpus for token warmup, instruction tuning, evaluation, and post-OCR text correction workflows. Dataset Summary The initial release consists of 4,087 verified entries targeting high-frequency and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-warmup-tokens.texttext-generation1K<n<10K0 likes12 downloads3mo agoHugging Face27johnzqlu /openr1_math_filtered_soft_tokens_en0.8_step3_tk5_tp1text10K<n<100K0 likes11 downloads8mo agoHugging Face28CL-From-Nothing /rlve-multitask-qwen3-4b-rollouts-n4-tokens16384tabular1K<n<10K0 likes10 downloads5mo agoHugging Face29procmarco /fpabl1-arm-a-web-tokens-48k fpabl1-arm-a-web-tokens-48k Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-a-web. Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121). Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total. Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks Source token composition: fw2_ita: 2,000,000,000… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-a-web-tokens-48k.tabularn<1K0 likes10 downloads3mo agoHugging Face30AiAF /TM_plain_qa_list_with_special_tokens.jsonltextn<1K0 likes9 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.