datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Qwen3.8-27B-Distill-1M-3.12B-Tokens
Qwen3.8-27B-Distill-1M-4.83B-Tokens
A unified, globally deduplicated, large-scale supervised distillation corpus built from 992,318 conversations generated by Qwen/Qwen3.8-27B, containing 4,834,771,862 target output tokens (3,570,459,498 reasoning tokens + 1,264,312,364 final response tokens) and 5,104,980,053 total sequence tokens.
1. Dataset Overview
This dataset merges, aligns, and deduplicates the two primary high-quality Qwen3.8-27B generation corpora on… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/Qwen3.8-27B-Distill-1M-3.12B-Tokens.fpabl1-arm-b-fp-tokens-48k
fpabl1-arm-b-fp-tokens-48k
Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-b-fp.
Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121).
Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total.
Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks
Source token composition:
fp_en: 1,000,000,000
fp_ita:… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-b-fp-tokens-48k.data-32k-200b-tokens
TR-HASH 32K · 200B Token Mixture
Pretokenized training mixture for compact TR-HASH language-model research.
The Dataset Viewer displays one summary row per source. The actual training data
is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin.
Each sequence contains 1,024 token IDs produced by the project 32K tokenizer.
Mixture
Source
Weight
Training tokens
DCLM
45%
90B
FineWeb-Edu deduplicated
30%
60B
Stack-Edu
10%
20B… See the full description on the dataset page: https://huggingface.co/datasets/Pacific-i64/data-32k-200b-tokens.RAW20seul_2316273_tokensRAW10seul_1262172_tokensmath-soft-tokens
Math Soft Tokens Dataset
Contains training steps: numinamath15_step_11_fixed.
French-Expert-SFT-81M-Tokens
French Expert SFT Corpus (81M Tokens)
🎯 Description
Ce dataset est un corpus de haute qualité conçu pour le Supervised Fine-Tuning (SFT). Il a été constitué par un moteur de recherche thématique profond (deep-crawl) ciblant les domaines de haute expertise technique et juridique française.
📊 Statistiques Clés
Nombre total de pépites (Samples) : 456,863
Volume estimé : ~81 Millions de Tokens
Taille moyenne par entrée : 629 caractères
Qualité : 0% doublons… See the full description on the dataset page: https://huggingface.co/datasets/Data-Elite/French-Expert-SFT-81M-Tokens.unnatural_code_instructions_20M_tokens_separatechinese-novel-continuation-precise-tokens
中文小说续写精确 Token 长度数据集
基于金庸《神雕侠侣》的高精度中文文本生成数据集,专为 GPU 内存测试和序列长度性能分析而设计。
数据集特点
高精度:99.4%+ 的目标 token 长度准确率
多种长度:5 个变体(1024、2048、4096、8192、16384 tokens)
统一格式:Alpaca 格式的小说续写任务
质量控制:95%+ 样本在目标长度的 ±2% 范围内
数据集统计
目标 Tokens
实际平均
准确率
样本数
文件大小
±1% 内样本
±2% 内样本
1024
1017.9
99.4%
800
2.5MB
739
775
2048
2037.7
99.5%
800
4.9MB
764
795
4096
4078.1
99.6%
800
9.8MB
792
800
8192
8158.1
99.6%
800
19.5MB
800
800
16384
16317.6
99.6%
800
39.1MB
800
800… See the full description on the dataset page: https://huggingface.co/datasets/aweffr/chinese-novel-continuation-precise-tokens.data-32k-200b-tokens
TR-HASH 32K · 200B Token Mixture
Pretokenized training mixture for compact TR-HASH language-model research.
The Dataset Viewer displays one summary row per source. The actual training data
is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin.
Each sequence contains 1,024 token IDs produced by the project 32K tokenizer.
Mixture
Source
Weight
Training tokens
DCLM
45%
90B
FineWeb-Edu deduplicated
30%
60B
Stack-Edu
10%
20B… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/data-32k-200b-tokens.pedro-open-dataset-max-512-tokenspedro-open-dataset-max-512-tokens-25kRAW15seul_1834048_tokensmodel_20_tokens_10_specsmodel_20_tokens_3_specsCritical-Tokens-Matter-Train-Datamodel_20_tokens_50_specspedro-open-dataset-max-512-tokens-10kInstruct-Data-7M-Tokensmodel_20_tokens_100_specsmodel_20_tokens_20_specsmodel_20_tokens_200_specsmodel_20_tokens_5_specsopencodegeneticinstruct-max-512-tokens-10ktransmla_pretrain_100m_tokenspashto-warmup-tokens
Pashto Warmup Tokens Dataset
This dataset contains a curated, deduplicated collection of high-quality, contextually accurate Pashto linguistic examples. It maps structural language tasks directly to the most critical vocabulary tokens in Pashto, providing a reliable corpus for token warmup, instruction tuning, evaluation, and post-OCR text correction workflows.
Dataset Summary
The initial release consists of 4,087 verified entries targeting high-frequency and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-warmup-tokens.openr1_math_filtered_soft_tokens_en0.8_step3_tk5_tp1rlve-multitask-qwen3-4b-rollouts-n4-tokens16384fpabl1-arm-a-web-tokens-48k
fpabl1-arm-a-web-tokens-48k
Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-a-web.
Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121).
Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total.
Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks
Source token composition:
fw2_ita: 2,000,000,000… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-a-web-tokens-48k.TM_plain_qa_list_with_special_tokens.jsonl
