CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-community /token-counts Marin Token Counts Token counts for all datasets used in Marin pretraining runs. Schema Column Type Description dataset string Dataset identifier marin_tokens int Number of tokens after tokenization category string Content domain (web, code, math, academic, books, etc.) synthetic bool Whether the data is LLM-generated or LLM-translated Categories web — Quality-classified Common Crawl text (Nemotron-CC) code — Source code and… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/token-counts.texttext-generationn<1K1 likes1.3k downloads27d agoHugging Face02Aipresso /prompts_under_512_tokens Under 512 Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use Dataset Overview Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles. 📊 Dataset Statistics Metric Value Total Files 200 Rows Per File 10,000 Total Rows 2,000,000 Token Range 1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.texttext-generation1M<n<10M0 likes103 downloads11mo agoHugging Face03Aipresso /medium_512_1k_tokens_prompts Medium 512-1K Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ By using this dataset you agree to our Terms of Use. Overview 703 high-quality English prompts whose length lies between 512 and 1 000 tokens.Every prompt has been de-duplicated, cleaned and token-counted with the GPT-2 tokenizer. Statistics Rows Token range File size Format 703 512 – 1 000 2.9 MB CSV Use-cases Medium-context language-model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/medium_512_1k_tokens_prompts.texttext-generationn<1K0 likes44 downloads11mo agoHugging Face04Aipresso /long_over_1k_tokens_prompts Long Over 1K Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use Dataset Overview Specialized collection of long-form English prompts (≥ 1 000 tokens) for training advanced models that require extensive context and complex reasoning. 📊 Dataset Statistics Metric Value Total Rows 289 Token Range 1 001 – 10 000 tokens File Size ≈ 3.7 MB Format Single CSV file Target… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/long_over_1k_tokens_prompts.texttext-generationn<1K0 likes39 downloads11mo agoHugging Face05AmanPriyanshu /Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens Dynamic Topic Modeling Dataset: RedPajama-1T SubSample (100k samples, 1k tokens) 📝Check out the Blog Post This dataset represents a curated subset of the RedPajama-1T Sample dataset, specifically processed for dynamic topic modeling applications. It contains 100,000 samples from the original dataset, with each document limited to the first 1,024 tokens for consistent processing. Dataset Overview Name:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens.textsummarization100K<n<1M8 likes28 downloads2y agoHugging Face06hadeelbkh /tokenized-IELTS-writing-task-2-evaluation-DialoGPT-mediumtexttext-generation1K<n<10K2 likes26 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.