CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MaxDevv /Qwen3.8-27B-Distill-1M-3.12B-Tokens Qwen3.8-27B-Distill-1M-4.83B-Tokens A unified, globally deduplicated, large-scale supervised distillation corpus built from 992,318 conversations generated by Qwen/Qwen3.8-27B, containing 4,834,771,862 target output tokens (3,570,459,498 reasoning tokens + 1,264,312,364 final response tokens) and 5,104,980,053 total sequence tokens. 1. Dataset Overview This dataset merges, aligns, and deduplicates the two primary high-quality Qwen3.8-27B generation corpora on… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/Qwen3.8-27B-Distill-1M-3.12B-Tokens.texttext-generation100K<n<1M1 likes316 downloads27d agoHugging Face02wmatejuk /midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to its EPR Labs source dataset. Based on MIDI datasets gathered by EPR Labs. Codec name: dyadic tokenizer vocab size: 512 max_time_step: 1.0 n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs.tabulartext-generation1M<n<10M0 likes229 downloads19d agoHugging Face03bobboyms /subset-Itau-Unibanco-aroeira-4B-tokens Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR) Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR) subset-Itau-Unibanco-aroeira-1B-tokens texttext-generation10M<n<100M1 likes157 downloads1y agoHugging Face04Pacific-i64 /data-32k-200b-tokens TR-HASH 32K · 200B Token Mixture Pretokenized training mixture for compact TR-HASH language-model research. The Dataset Viewer displays one summary row per source. The actual training data is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin. Each sequence contains 1,024 token IDs produced by the project 32K tokenizer. Mixture Source Weight Training tokens DCLM 45% 90B FineWeb-Edu deduplicated 30% 60B Stack-Edu 10% 20B… See the full description on the dataset page: https://huggingface.co/datasets/Pacific-i64/data-32k-200b-tokens.tabulartext-generationn<1K1 likes114 downloads1mo agoHugging Face05Aipresso /prompts_under_512_tokens Under 512 Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use Dataset Overview Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles. 📊 Dataset Statistics Metric Value Total Files 200 Rows Per File 10,000 Total Rows 2,000,000 Token Range 1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.texttext-generation1M<n<10M0 likes107 downloads11mo agoHugging Face06marin-community /open-thoughts-4-30k-code-qwen3-32b-annotated-32768-tokens Dataset Card for Open-Thoughts-4-30K-Code-Qwen3-32B-Annotated-32768-Tokens Overview This dataset is a variant of marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated with an extended maximum sequence length. The responses in the generated_text column were generated with max output tokens = 32768 (instead of 7500 in the original dataset), allowing for longer and more complete chain-of-thought reasoning. Generation Details Model: Qwen/Qwen3-32B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated-32768-tokens.tabulartext-generation10K<n<100K0 likes107 downloads8mo agoHugging Face07wmatejuk /midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512 midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512 Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to Maestro, GiantMIDI, ATEPP, or MusicNet. Based on MIDI datasets gathered by EPR Labs. Codec name: dyadic tokenizer vocab size: 512 max_time_step: 1.0 n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512.tabulartext-generation10K<n<100K0 likes70 downloads1mo agoHugging Face08Ranjit89 /Assamese-Text-Dataset-45T-Tokens I have massive Assamese Dataset nearly about 45.3T (45333004592600) Tokens It has a lots of Assamese sentances from various sources, 99.9999% of the dataset are cleanned just download the backup_data.tar.zst file and start using it. happy training.... My email: ranjitdax89@gmail.com At least share your opinion… or maybe a simple “thanks” 😄 Topic / Dataset Tokens Approx. Scale Source Poems Dataset 92.6K 0.0000926B… See the full description on the dataset page: https://huggingface.co/datasets/Ranjit89/Assamese-Text-Dataset-45T-Tokens.texttext-generationn<1K0 likes54 downloads4mo agoHugging Face09spadeMIA /pmc_finetune_corpus_1024-2040_tokens PMC 1024-2040 Biomedical Fine-Tuning Corpus Summary This is a cleaned biomedical long-text corpus for autoregressive language-model fine-tuning and held-out evaluation. split rows role train 10,000 fine-tuning test 1,000 held-out evaluation The public schema is text-only: text: string No PMCID, date, license, URL, or provenance fields are included in the public dataset files. Token Contract The corpus is built for… See the full description on the dataset page: https://huggingface.co/datasets/spadeMIA/pmc_finetune_corpus_1024-2040_tokens.texttext-generation10K<n<100K0 likes49 downloads2mo agoHugging Face10Aipresso /medium_512_1k_tokens_prompts Medium 512-1K Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ By using this dataset you agree to our Terms of Use. Overview 703 high-quality English prompts whose length lies between 512 and 1 000 tokens.Every prompt has been de-duplicated, cleaned and token-counted with the GPT-2 tokenizer. Statistics Rows Token range File size Format 703 512 – 1 000 2.9 MB CSV Use-cases Medium-context language-model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/medium_512_1k_tokens_prompts.texttext-generationn<1K0 likes42 downloads11mo agoHugging Face11AETHORIA-AI /data-32k-200b-tokens TR-HASH 32K · 200B Token Mixture Pretokenized training mixture for compact TR-HASH language-model research. The Dataset Viewer displays one summary row per source. The actual training data is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin. Each sequence contains 1,024 token IDs produced by the project 32K tokenizer. Mixture Source Weight Training tokens DCLM 45% 90B FineWeb-Edu deduplicated 30% 60B Stack-Edu 10% 20B… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/data-32k-200b-tokens.tabulartext-generationn<1K0 likes40 downloads1mo agoHugging Face12Aipresso /long_over_1k_tokens_prompts Long Over 1K Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use Dataset Overview Specialized collection of long-form English prompts (≥ 1 000 tokens) for training advanced models that require extensive context and complex reasoning. 📊 Dataset Statistics Metric Value Total Rows 289 Token Range 1 001 – 10 000 tokens File Size ≈ 3.7 MB Format Single CSV file Target… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/long_over_1k_tokens_prompts.texttext-generationn<1K0 likes39 downloads11mo agoHugging Face13pedrodev2026 /pedro-open-dataset-max-512-tokenstexttext-generation10K<n<100K0 likes35 downloads7mo agoHugging Face14pedrodev2026 /pedro-open-dataset-max-512-tokens-25ktexttext-generation10K<n<100K0 likes35 downloads7mo agoHugging Face15AmanPriyanshu /Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens Dynamic Topic Modeling Dataset: RedPajama-1T SubSample (100k samples, 1k tokens) 📝Check out the Blog Post This dataset represents a curated subset of the RedPajama-1T Sample dataset, specifically processed for dynamic topic modeling applications. It contains 100,000 samples from the original dataset, with each document limited to the first 1,024 tokens for consistent processing. Dataset Overview Name:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens.textsummarization100K<n<1M8 likes28 downloads2y agoHugging Face16pedrodev2026 /pedro-open-dataset-max-512-tokens-10ktexttext-generation10K<n<100K0 likes23 downloads7mo agoHugging Face17bobboyms /subset-Itau-Unibanco-aroeira-1B-tokens Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR) Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR) subset-Itau-Unibanco-aroeira-1B-tokens texttext-generation1M<n<10M2 likes22 downloads1y agoHugging Face18nassimjp /pashto-warmup-tokens Pashto Warmup Tokens Dataset This dataset contains a curated, deduplicated collection of high-quality, contextually accurate Pashto linguistic examples. It maps structural language tasks directly to the most critical vocabulary tokens in Pashto, providing a reliable corpus for token warmup, instruction tuning, evaluation, and post-OCR text correction workflows. Dataset Summary The initial release consists of 4,087 verified entries targeting high-frequency and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-warmup-tokens.texttext-generation1K<n<10K0 likes12 downloads3mo agoHugging Face19Omarrran /KS-PRET-5M_5_million_kashmiri_Pretrainning_LLM_dataset_12M_tokens_2026gated KS-PRET-5M: Kashmiri Pretraining Corpus 5,090,244 words · ~12,130,000 subword tokens · 295,433 vocabulary · April 2026 Building the complete AI stack for Kashmiri — 7 million speakers, virtually no prior computational resources. Dataset Summary KS-PRET-5M is a large-scale, deeply cleaned Kashmiri language pretraining corpus — the largest publicly available dataset for the Kashmiri language. All text is formatted as a single continuous stream, the standard… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/KS-PRET-5M_5_million_kashmiri_Pretrainning_LLM_dataset_12M_tokens_2026.texttext-generationn<1K1 likes10 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.