datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fpabl1-arm-b-fp-tokens-48k
fpabl1-arm-b-fp-tokens-48k
Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-b-fp.
Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121).
Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total.
Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks
Source token composition:
fp_en: 1,000,000,000
fp_ita:… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-b-fp-tokens-48k.data-32k-200b-tokens
TR-HASH 32K · 200B Token Mixture
Pretokenized training mixture for compact TR-HASH language-model research.
The Dataset Viewer displays one summary row per source. The actual training data
is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin.
Each sequence contains 1,024 token IDs produced by the project 32K tokenizer.
Mixture
Source
Weight
Training tokens
DCLM
45%
90B
FineWeb-Edu deduplicated
30%
60B
Stack-Edu
10%
20B… See the full description on the dataset page: https://huggingface.co/datasets/Pacific-i64/data-32k-200b-tokens.data-32k-200b-tokens
TR-HASH 32K · 200B Token Mixture
Pretokenized training mixture for compact TR-HASH language-model research.
The Dataset Viewer displays one summary row per source. The actual training data
is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin.
Each sequence contains 1,024 token IDs produced by the project 32K tokenizer.
Mixture
Source
Weight
Training tokens
DCLM
45%
90B
FineWeb-Edu deduplicated
30%
60B
Stack-Edu
10%
20B… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/data-32k-200b-tokens.model_20_tokens_10_specsmodel_20_tokens_3_specsmodel_20_tokens_50_specsmodel_20_tokens_100_specsmodel_20_tokens_20_specsmodel_20_tokens_200_specsmodel_20_tokens_5_specsrlve-multitask-qwen3-4b-rollouts-n4-tokens16384fpabl1-arm-a-web-tokens-48k
fpabl1-arm-a-web-tokens-48k
Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-a-web.
Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121).
Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total.
Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks
Source token composition:
fw2_ita: 2,000,000,000… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-a-web-tokens-48k.tokenskipexpresso-conversational-tokens
MintTTS Pre-tokenized Audio Tokens
Pre-extracted audio codec tokens for TTS training.
Source
Dataset: somu9/expresso-conversational
Codec: MOSS-Audio-Tokenizer-Nano
Codec sample rate: 48,000 Hz (stereo)
Frame rate: 12.5 Hz (1 frame = 80ms)
Stats
Metric
Value
Total samples
29,487
Total audio hours
27.8h
Codebooks
16
Avg frames/sample
42.5
Avg duration
3.4s
Format
JSONL file (manifest.jsonl) where each line is:
{
"text":… See the full description on the dataset page: https://huggingface.co/datasets/somu9/expresso-conversational-tokens.nonverbal-tts-filtered-tokens
MintTTS Pre-tokenized Audio Tokens
Pre-extracted audio codec tokens for TTS training.
Source
Dataset: somu9/nonverbal-tts-filtered
Codec: MOSS-Audio-Tokenizer-Nano
Codec sample rate: 48,000 Hz (stereo)
Frame rate: 12.5 Hz (1 frame = 80ms)
Stats
Metric
Value
Total samples
6,082
Total audio hours
15.6h
Codebooks
16
Avg frames/sample
115.3
Avg duration
9.2s
Format
JSONL file (manifest.jsonl) where each line is:
{
"text": "The… See the full description on the dataset page: https://huggingface.co/datasets/somu9/nonverbal-tts-filtered-tokens.minesweeper-student-kukurasu20k-qwen1.7b-e3-mask-rollouts-n2-tokens16384mls10k-tokens
MintTTS Pre-tokenized Audio Tokens
Pre-extracted audio codec tokens for TTS training.
Source
Dataset: parler-tts/mls_eng_10k
Codec: MOSS-Audio-Tokenizer-Nano
Codec sample rate: 48,000 Hz (stereo)
Frame rate: 12.5 Hz (1 frame = 80ms)
Stats
Metric
Value
Total samples
2,420,047
Total audio hours
9,976.5h
Codebooks
16
Avg frames/sample
185.5
Avg duration
14.8s
Format
JSONL file (manifest.jsonl) where each line is:
{
"text":… See the full description on the dataset page: https://huggingface.co/datasets/somu9/mls10k-tokens.mls_eng_tokens
MLS English - Pre-tokenized Audio Codec Tokens
Pre-extracted audio codec tokens from the Multilingual LibriSpeech (MLS) English dataset, tokenized using MOSS-Audio-Tokenizer for text-to-speech training.
Source
Property
Value
Source dataset
parler-tts/mls_eng
Audio codec
MOSS-Audio-Tokenizer
Language
English
Codec sample rate
48,000 Hz (stereo)
Frame rate
12.5 Hz (1 frame = 80ms)
Codebooks
16 (RVQ, 1024 vocab each)
Splits included
train, dev, test… See the full description on the dataset page: https://huggingface.co/datasets/somu9/mls_eng_tokens.hindi-hq-tokens
MintTTS Pre-tokenized Audio Tokens
Pre-extracted audio codec tokens for TTS training.
Source
Dataset: somu9/hindi-hq
Codec: MOSS-Audio-Tokenizer-Nano
Codec sample rate: 48,000 Hz (stereo)
Frame rate: 12.5 Hz (1 frame = 80ms)
Stats
Metric
Value
Total samples
439,507
Total audio hours
811.7h
Codebooks
16
Avg frames/sample
83.1
Avg duration
6.6s
Format
JSONL file (manifest.jsonl) where each line is:
{
"text":… See the full description on the dataset page: https://huggingface.co/datasets/somu9/hindi-hq-tokens.250-max-tokens-trash-15maskminesweeper-teacher-kukurasu20k-qwen1.7b-e3-mask-rollouts-n2-tokens16384libritts-clean-v2-tokens
MintTTS Pre-tokenized Audio Tokens
Pre-extracted audio codec tokens for TTS training.
Source
Dataset: mythicinfinity/libritts_r
Codec: MOSS-Audio-Tokenizer-Nano
Codec sample rate: 48,000 Hz (stereo)
Frame rate: 12.5 Hz (1 frame = 80ms)
Stats
Metric
Value
Total samples
148,954
Total audio hours
242.6h
Codebooks
16
Avg frames/sample
73.3
Avg duration
5.9s
Format
JSONL file (manifest.jsonl) where each line is:
{
"text": "The… See the full description on the dataset page: https://huggingface.co/datasets/somu9/libritts-clean-v2-tokens.expresso-tokens
MintTTS Pre-tokenized Audio Tokens
Pre-extracted audio codec tokens for TTS training.
Source
Dataset: ylacombe/expresso
Codec: MOSS-Audio-Tokenizer-Nano
Codec sample rate: 48,000 Hz (stereo)
Frame rate: 12.5 Hz (1 frame = 80ms)
Stats
Metric
Value
Total samples
11,614
Total audio hours
10.8h
Codebooks
16
Avg frames/sample
41.8
Avg duration
3.3s
Format
JSONL file (manifest.jsonl) where each line is:
{
"text": "The… See the full description on the dataset page: https://huggingface.co/datasets/somu9/expresso-tokens.20k-max-tokens-022125-step0-aug1-ttt3-0-batch020k-max-tokens-022125-step0-aug2-ttt6-0-batch020k-max-tokens-022125-step0-aug1-ttt4-0-batch220k-max-tokens-022125-step0-aug2-ttt6-0-batch420k-max-tokens-022125-step0-aug0-ttt2-0-batch5sudoku-student-minekuk-nemtron8b-rollouts-n2-tokens16384jenny_30h-tokens
MintTTS Pre-tokenized Audio Tokens
Pre-extracted audio codec tokens for TTS training.
Source
Dataset: reach-vb/jenny_tts_dataset
Codec: MOSS-Audio-Tokenizer-Nano
Codec sample rate: 48,000 Hz (stereo)
Frame rate: 12.5 Hz (1 frame = 80ms)
Stats
Metric
Value
Total samples
20,141
Total audio hours
26.4h
Codebooks
16
Avg frames/sample
59.0
Avg duration
4.7s
Format
JSONL file (manifest.jsonl) where each line is:
{
"text": "The… See the full description on the dataset page: https://huggingface.co/datasets/somu9/jenny_30h-tokens.
