CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-community /token-counts Marin Token Counts Token counts for all datasets used in Marin pretraining runs. Schema Column Type Description dataset string Dataset identifier marin_tokens int Number of tokens after tokenization category string Content domain (web, code, math, academic, books, etc.) synthetic bool Whether the data is LLM-generated or LLM-translated Categories web — Quality-classified Common Crawl text (Nemotron-CC) code — Source code and… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/token-counts.texttext-generationn<1K1 likes1.3k downloads26d agoHugging Face02blanchon /parler-tts_mls_eng_10k_snac_token_old Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tabularautomatic-speech-recognition100K<n<1M1 likes991 downloads2y agoHugging Face03SakethVemula /fixed-tokenizer-morphscore-segmentstabular10M<n<100M0 likes534 downloads6mo agoHugging Face04nancyH /ablation_tokenstext1M<n<10M0 likes152 downloads6mo agoHugging Face05Aipresso /prompts_under_512_tokens Under 512 Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use Dataset Overview Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles. 📊 Dataset Statistics Metric Value Total Files 200 Rows Per File 10,000 Total Rows 2,000,000 Token Range 1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.texttext-generation1M<n<10M0 likes99 downloads11mo agoHugging Face06globalise /globalise_NER_token_classification_dataset Dataset Card for Dataset Name The globalise_NER_token_classification dataset is a fine-grained dataset for the training of token-classification NER models on Dutch East-India Company texts (17th to 18th century). Dataset Details Dataset Description The dataset provides 15 fine-grained labels detailing activities and people of the Dutch East-India Company (VOC), and can be used to train NER token-classification models for the period 17th-18th century and the… See the full description on the dataset page: https://huggingface.co/datasets/globalise/globalise_NER_token_classification_dataset.texttoken-classificationn<1K1 likes64 downloads11mo agoHugging Face07Lyte /tokenizer-leaderboard Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): en License: mit Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use [More… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/tokenizer-leaderboard.tabularn<1K0 likes60 downloads4mo agoHugging Face08Adapting /empathetic_dialogues_with_special_tokenstabular10K<n<100K2 likes55 downloads4y agoHugging Face09anismahmahi /classification_token_propagandatexttoken-classificationn<1K0 likes55 downloads2y agoHugging Face10nishita /webnlg_tokenstext1K<n<10K0 likes48 downloads4y agoHugging Face11Aipresso /medium_512_1k_tokens_prompts Medium 512-1K Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ By using this dataset you agree to our Terms of Use. Overview 703 high-quality English prompts whose length lies between 512 and 1 000 tokens.Every prompt has been de-duplicated, cleaned and token-counted with the GPT-2 tokenizer. Statistics Rows Token range File size Format 703 512 – 1 000 2.9 MB CSV Use-cases Medium-context language-model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/medium_512_1k_tokens_prompts.texttext-generationn<1K0 likes43 downloads11mo agoHugging Face12LabARSS /MMLU-Pro-single-token-entropy Dataset Card for MMLU Pro with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B MMLU Pro dataset with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B Dataset Details Dataset Description Following up on the results from "When an LLM is apprehensive about its answers -- and when its uncertainty is justified", we measure the response entopy for MMLU Pro dataset when the model is prompted to… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-single-token-entropy.tabular100K<n<1M0 likes42 downloads1y agoHugging Face13Aipresso /long_over_1k_tokens_prompts Long Over 1K Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use Dataset Overview Specialized collection of long-form English prompts (≥ 1 000 tokens) for training advanced models that require extensive context and complex reasoning. 📊 Dataset Statistics Metric Value Total Rows 289 Token Range 1 001 – 10 000 tokens File Size ≈ 3.7 MB Format Single CSV file Target… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/long_over_1k_tokens_prompts.texttext-generationn<1K0 likes40 downloads11mo agoHugging Face14alxfgh /PubChem10M_SELFIES_TokenizedCustom cl100k tokenized version of PubChem10M_SELFIES. text1M<n<10M2 likes33 downloads3y agoHugging Face15TokenfreeEMNLPSubmission /SimpleDC simpledc-dataset Official huggingface dataset for the SimpleDC (Simple Digestive Cancer) dataset Please cite as: @article{rahman2024health, title={Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning}, author={Rahman, Md Mushfiqur and Irbaz, Mohammad Sabik and North, Kai and Williams, Michelle S and Zampieri, Marcos and Lybarger, Kevin}, journal={arXiv preprint arXiv:2401.15043}, year={2024} } text1K<n<10K0 likes33 downloads2y agoHugging Face16hadeelbkh /tokenized-IELTS-writing-task-2-evaluation-DialoGPT-mediumtexttext-generation1K<n<10K2 likes30 downloads1y agoHugging Face17token-opt-org /Token_Optimization_Org AI Safety & Bias Evaluation Conversations Dataset Summary This dataset contains simulated multi-turn conversations designed to evaluate AI language model behavior across two safety-critical domains: self-harm response handling and political bias. Each row represents a single evaluation scenario where an AI model's responses are assessed for safety compliance or neutrality. The dataset is intended to support research and development of safer, less biased AI systems. All… See the full description on the dataset page: https://huggingface.co/datasets/token-opt-org/Token_Optimization_Org.texttext-classificationn<1K0 likes30 downloads4mo agoHugging Face18sajjadanwar0 /token-budgets-catalog Token Budgets — Empirical catalogue and inter-rater reliability data Data for: Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun Incidents, with an Affine-Typed Rust Mitigation as a Case Study. Sajjad Khan, 2026. arXiv:2606.04056 — preprint. This dataset bundles three artefacts referenced in the paper: catalogue — the harvested catalogue of LLM-agent budget-overrun incidents across 21 orchestration frameworks (2023–2026), 167 rows total. The IRR-included… See the full description on the dataset page: https://huggingface.co/datasets/sajjadanwar0/token-budgets-catalog.texttext-classificationn<1K0 likes29 downloads4mo agoHugging Face19AmanPriyanshu /Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens Dynamic Topic Modeling Dataset: RedPajama-1T SubSample (100k samples, 1k tokens) 📝Check out the Blog Post This dataset represents a curated subset of the RedPajama-1T Sample dataset, specifically processed for dynamic topic modeling applications. It contains 100,000 samples from the original dataset, with each document limited to the first 1,024 tokens for consistent processing. Dataset Overview Name:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens.textsummarization100K<n<1M8 likes28 downloads2y agoHugging Face20distil-whisper /whisper_transcriptions_token_idstext100K<n<1M0 likes27 downloads3y agoHugging Face21Circularmachines /Batch_indexing_machine_tokenstabular1M<n<10M0 likes22 downloads3y agoHugging Face22NhungNguyen /rams-no-special-tokenstext1K<n<10K0 likes21 downloads1y agoHugging Face23ClarusC64 /clinical-healing-trajectory-tokenization-phase-segmentation-v0.1What this dataset tests Whether a model can segment high-frequency recovery datainto interpretable healing phases. Required outputs phase_sequence phase_boundaries phase_confidence_0_100 Token labels acute_drop early_rebound consolidation_plateau oscillatory_instability secondary_drop delayed_rebound steady_ascent maladaptive_plateau recovery_lock_in Boundary format Use day indicesexampleacute_drop d0-d2 Typical failures naming phases without boundaries… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-healing-trajectory-tokenization-phase-segmentation-v0.1.tabulartext-classificationn<1K0 likes19 downloads8mo agoHugging Face24TokenBender /sentence_retrieval_hindi_SFTtext10K<n<100K2 likes18 downloads3y agoHugging Face25TokenBender /e5_FT_sentence_retrieval_task_Hinditext10K<n<100K0 likes17 downloads3y agoHugging Face26maneln /tokenized_datasetQAtextn<1K0 likes16 downloads2y agoHugging Face27CentificAIResearch /token-optimization AI Safety & Bias Evaluation Conversations Dataset Summary This dataset contains simulated multi-turn conversations designed to evaluate AI language model behavior across two safety-critical domains: self-harm response handling and political bias. Each row represents a single evaluation scenario where an AI model's responses are assessed for safety compliance or neutrality. The dataset is intended to support research and development of safer, less biased AI… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/token-optimization.texttext-classificationn<1K1 likes14 downloads3mo agoHugging Face28AmanPriyanshu /GTE-ModernBERT-RedPajama-Data-1T-100k-SubSample-max-1k-tokenstext100K<n<1M0 likes13 downloads2y agoHugging Face29Gugu8 /Token-Efficiency token_efficiency_corpus A 2.5 GB CSV corpus teaching LLMs to minimize token usage in their outputs. Progresses from basic filler removal to expert-level nested reasoning compression. Contents verbose_output - The padded, wasteful version of the text efficient_output - The compressed, token-efficient equivalent technique - Compression strategy used subcategory - Specific variant of the technique difficulty - Tier 1 (easiest)… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Token-Efficiency.tabular1M<n<10M0 likes13 downloads2mo agoHugging Face30TokenBender /Hindi_SFT_sentence_retriever_settext10K<n<100K1 likes12 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.