CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01procmarco /fpabl1-arm-b-fp-tokens-48k fpabl1-arm-b-fp-tokens-48k Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-b-fp. Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121). Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total. Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks Source token composition: fp_en: 1,000,000,000 fp_ita:… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-b-fp-tokens-48k.tabularn<1K0 likes178 downloads3mo agoHugging Face02Pacific-i64 /data-32k-200b-tokens TR-HASH 32K · 200B Token Mixture Pretokenized training mixture for compact TR-HASH language-model research. The Dataset Viewer displays one summary row per source. The actual training data is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin. Each sequence contains 1,024 token IDs produced by the project 32K tokenizer. Mixture Source Weight Training tokens DCLM 45% 90B FineWeb-Edu deduplicated 30% 60B Stack-Edu 10% 20B… See the full description on the dataset page: https://huggingface.co/datasets/Pacific-i64/data-32k-200b-tokens.tabulartext-generationn<1K1 likes114 downloads1mo agoHugging Face03AETHORIA-AI /data-32k-200b-tokens TR-HASH 32K · 200B Token Mixture Pretokenized training mixture for compact TR-HASH language-model research. The Dataset Viewer displays one summary row per source. The actual training data is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin. Each sequence contains 1,024 token IDs produced by the project 32K tokenizer. Mixture Source Weight Training tokens DCLM 45% 90B FineWeb-Edu deduplicated 30% 60B Stack-Edu 10% 20B… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/data-32k-200b-tokens.tabulartext-generationn<1K0 likes40 downloads1mo agoHugging Face04cotarc0 /model_20_tokens_10_specstabularn<1K0 likes33 downloads1y agoHugging Face05cotarc0 /model_20_tokens_3_specstabularn<1K0 likes28 downloads1y agoHugging Face06cotarc0 /model_20_tokens_50_specstabularn<1K0 likes24 downloads1y agoHugging Face07cotarc0 /model_20_tokens_100_specstabularn<1K0 likes21 downloads1y agoHugging Face08cotarc0 /model_20_tokens_20_specstabularn<1K0 likes20 downloads1y agoHugging Face09cotarc0 /model_20_tokens_200_specstabularn<1K0 likes17 downloads1y agoHugging Face10cotarc0 /model_20_tokens_5_specstabularn<1K0 likes17 downloads1y agoHugging Face11CL-From-Nothing /rlve-multitask-qwen3-4b-rollouts-n4-tokens16384tabular1K<n<10K0 likes10 downloads5mo agoHugging Face12procmarco /fpabl1-arm-a-web-tokens-48k fpabl1-arm-a-web-tokens-48k Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-a-web. Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121). Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total. Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks Source token composition: fw2_ita: 2,000,000,000… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-a-web-tokens-48k.tabularn<1K0 likes10 downloads3mo agoHugging Face13yiboowang /tokenskiptabular1K<n<10K0 likes9 downloads1y agoHugging Face14somu9 /expresso-conversational-tokensgated MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: somu9/expresso-conversational Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 29,487 Total audio hours 27.8h Codebooks 16 Avg frames/sample 42.5 Avg duration 3.4s Format JSONL file (manifest.jsonl) where each line is: { "text":… See the full description on the dataset page: https://huggingface.co/datasets/somu9/expresso-conversational-tokens.tabulartext-to-speech10K<n<100K0 likes9 downloads4mo agoHugging Face15somu9 /nonverbal-tts-filtered-tokensgated MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: somu9/nonverbal-tts-filtered Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 6,082 Total audio hours 15.6h Codebooks 16 Avg frames/sample 115.3 Avg duration 9.2s Format JSONL file (manifest.jsonl) where each line is: { "text": "The… See the full description on the dataset page: https://huggingface.co/datasets/somu9/nonverbal-tts-filtered-tokens.tabulartext-to-speech1K<n<10K1 likes8 downloads4mo agoHugging Face16CL-From-Nothing /minesweeper-student-kukurasu20k-qwen1.7b-e3-mask-rollouts-n2-tokens16384tabular10K<n<100K0 likes7 downloads6mo agoHugging Face17somu9 /mls10k-tokensgated MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: parler-tts/mls_eng_10k Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 2,420,047 Total audio hours 9,976.5h Codebooks 16 Avg frames/sample 185.5 Avg duration 14.8s Format JSONL file (manifest.jsonl) where each line is: { "text":… See the full description on the dataset page: https://huggingface.co/datasets/somu9/mls10k-tokens.tabulartext-to-speech1M<n<10M0 likes7 downloads5mo agoHugging Face18somu9 /mls_eng_tokensgated MLS English - Pre-tokenized Audio Codec Tokens Pre-extracted audio codec tokens from the Multilingual LibriSpeech (MLS) English dataset, tokenized using MOSS-Audio-Tokenizer for text-to-speech training. Source Property Value Source dataset parler-tts/mls_eng Audio codec MOSS-Audio-Tokenizer Language English Codec sample rate 48,000 Hz (stereo) Frame rate 12.5 Hz (1 frame = 80ms) Codebooks 16 (RVQ, 1024 vocab each) Splits included train, dev, test… See the full description on the dataset page: https://huggingface.co/datasets/somu9/mls_eng_tokens.tabulartext-to-speech1M<n<10M1 likes7 downloads4mo agoHugging Face19somu9 /hindi-hq-tokensgated MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: somu9/hindi-hq Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 439,507 Total audio hours 811.7h Codebooks 16 Avg frames/sample 83.1 Avg duration 6.6s Format JSONL file (manifest.jsonl) where each line is: { "text":… See the full description on the dataset page: https://huggingface.co/datasets/somu9/hindi-hq-tokens.tabulartext-to-speech100K<n<1M1 likes7 downloads3mo agoHugging Face20tttx /250-max-tokens-trash-15masktabularn<1K0 likes6 downloads2y agoHugging Face21CL-From-Nothing /minesweeper-teacher-kukurasu20k-qwen1.7b-e3-mask-rollouts-n2-tokens16384tabular10K<n<100K0 likes6 downloads6mo agoHugging Face22somu9 /libritts-clean-v2-tokensgated MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: mythicinfinity/libritts_r Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 148,954 Total audio hours 242.6h Codebooks 16 Avg frames/sample 73.3 Avg duration 5.9s Format JSONL file (manifest.jsonl) where each line is: { "text": "The… See the full description on the dataset page: https://huggingface.co/datasets/somu9/libritts-clean-v2-tokens.tabulartext-to-speech100K<n<1M0 likes6 downloads5mo agoHugging Face23somu9 /expresso-tokensgated MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: ylacombe/expresso Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 11,614 Total audio hours 10.8h Codebooks 16 Avg frames/sample 41.8 Avg duration 3.3s Format JSONL file (manifest.jsonl) where each line is: { "text": "The… See the full description on the dataset page: https://huggingface.co/datasets/somu9/expresso-tokens.tabulartext-to-speech10K<n<100K0 likes5 downloads5mo agoHugging Face24tttx /20k-max-tokens-022125-step0-aug1-ttt3-0-batch0tabularn<1K0 likes4 downloads2y agoHugging Face25tttx /20k-max-tokens-022125-step0-aug2-ttt6-0-batch0tabularn<1K0 likes4 downloads2y agoHugging Face26tttx /20k-max-tokens-022125-step0-aug1-ttt4-0-batch2tabularn<1K0 likes4 downloads2y agoHugging Face27tttx /20k-max-tokens-022125-step0-aug2-ttt6-0-batch4tabularn<1K0 likes4 downloads2y agoHugging Face28tttx /20k-max-tokens-022125-step0-aug0-ttt2-0-batch5tabularn<1K0 likes4 downloads2y agoHugging Face29CL-From-Nothing /sudoku-student-minekuk-nemtron8b-rollouts-n2-tokens16384tabular10K<n<100K0 likes4 downloads6mo agoHugging Face30somu9 /jenny_30h-tokensgated MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: reach-vb/jenny_tts_dataset Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 20,141 Total audio hours 26.4h Codebooks 16 Avg frames/sample 59.0 Avg duration 4.7s Format JSONL file (manifest.jsonl) where each line is: { "text": "The… See the full description on the dataset page: https://huggingface.co/datasets/somu9/jenny_30h-tokens.tabulartext-to-speech10K<n<100K1 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.