CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-community /token-counts Marin Token Counts Token counts for all datasets used in Marin pretraining runs. Schema Column Type Description dataset string Dataset identifier marin_tokens int Number of tokens after tokenization category string Content domain (web, code, math, academic, books, etc.) synthetic bool Whether the data is LLM-generated or LLM-translated Categories web — Quality-classified Common Crawl text (Nemotron-CC) code — Source code and… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/token-counts.texttext-generationn<1K1 likes1.3k downloads29d agoHugging Face02blanchon /parler-tts_mls_eng_10k_snac_token_old Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tabularautomatic-speech-recognition100K<n<1M1 likes998 downloads2y agoHugging Face03SakethVemula /fixed-tokenizer-morphscore-segmentstabular10M<n<100M0 likes534 downloads6mo agoHugging Face04nancyH /ablation_tokenstext1M<n<10M0 likes118 downloads6mo agoHugging Face05Aipresso /prompts_under_512_tokens Under 512 Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use Dataset Overview Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles. 📊 Dataset Statistics Metric Value Total Files 200 Rows Per File 10,000 Total Rows 2,000,000 Token Range 1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.texttext-generation1M<n<10M0 likes106 downloads11mo agoHugging Face06globalise /globalise_NER_token_classification_dataset Dataset Card for Dataset Name The globalise_NER_token_classification dataset is a fine-grained dataset for the training of token-classification NER models on Dutch East-India Company texts (17th to 18th century). Dataset Details Dataset Description The dataset provides 15 fine-grained labels detailing activities and people of the Dutch East-India Company (VOC), and can be used to train NER token-classification models for the period 17th-18th century and the… See the full description on the dataset page: https://huggingface.co/datasets/globalise/globalise_NER_token_classification_dataset.texttoken-classificationn<1K1 likes63 downloads1y agoHugging Face07tokeron /Piyyuttabulartext-classification10K<n<100K0 likes61 downloads3y agoHugging Face08Lyte /tokenizer-leaderboard Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): en License: mit Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use [More… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/tokenizer-leaderboard.tabularn<1K0 likes58 downloads4mo agoHugging Face09anismahmahi /classification_token_propagandatexttoken-classificationn<1K0 likes56 downloads2y agoHugging Face10Adapting /empathetic_dialogues_with_special_tokenstabular10K<n<100K2 likes55 downloads4y agoHugging Face11farhamu /tokopedia-product-reviews-2019 Tokopedia Product Reviews 2019 Dataset Description This dataset contains 40,607 product reviews from Tokopedia, one of Indonesia's largest e-commerce platforms, scraped in 2019. The dataset provides valuable insights into customer sentiment and shopping behavior in the Indonesian e-commerce market. Dataset Summary Language: Indonesian (Bahasa Indonesia) Task: Sentiment Analysis, Product Review Analysis, E-commerce Research Size: 40,607 reviews Categories: 5… See the full description on the dataset page: https://huggingface.co/datasets/farhamu/tokopedia-product-reviews-2019.tabular10K<n<100K1 likes54 downloads1y agoHugging Face12Tokyomonster /JBB-Behaviors An Open Robustness Benchmark for Jailbreaking Language Models NeurIPS 2024 Datasets and Benchmarks Track Paper | Leaderboard | Benchmark code What is JailbreakBench? Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/Tokyomonster/JBB-Behaviors.tabularn<1K0 likes51 downloads6mo agoHugging Face13nishita /webnlg_tokenstext1K<n<10K0 likes44 downloads4y agoHugging Face14Aipresso /medium_512_1k_tokens_prompts Medium 512-1K Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ By using this dataset you agree to our Terms of Use. Overview 703 high-quality English prompts whose length lies between 512 and 1 000 tokens.Every prompt has been de-duplicated, cleaned and token-counted with the GPT-2 tokenizer. Statistics Rows Token range File size Format 703 512 – 1 000 2.9 MB CSV Use-cases Medium-context language-model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/medium_512_1k_tokens_prompts.texttext-generationn<1K0 likes44 downloads11mo agoHugging Face15LabARSS /MMLU-Pro-single-token-entropy Dataset Card for MMLU Pro with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B MMLU Pro dataset with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B Dataset Details Dataset Description Following up on the results from "When an LLM is apprehensive about its answers -- and when its uncertainty is justified", we measure the response entopy for MMLU Pro dataset when the model is prompted to… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-single-token-entropy.tabular100K<n<1M0 likes42 downloads1y agoHugging Face16Aipresso /long_over_1k_tokens_prompts Long Over 1K Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use Dataset Overview Specialized collection of long-form English prompts (≥ 1 000 tokens) for training advanced models that require extensive context and complex reasoning. 📊 Dataset Statistics Metric Value Total Rows 289 Token Range 1 001 – 10 000 tokens File Size ≈ 3.7 MB Format Single CSV file Target… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/long_over_1k_tokens_prompts.texttext-generationn<1K0 likes39 downloads11mo agoHugging Face17TokenfreeEMNLPSubmission /SimpleDC simpledc-dataset Official huggingface dataset for the SimpleDC (Simple Digestive Cancer) dataset Please cite as: @article{rahman2024health, title={Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning}, author={Rahman, Md Mushfiqur and Irbaz, Mohammad Sabik and North, Kai and Williams, Michelle S and Zampieri, Marcos and Lybarger, Kevin}, journal={arXiv preprint arXiv:2401.15043}, year={2024} } text1K<n<10K0 likes38 downloads2y agoHugging Face18alxfgh /PubChem10M_SELFIES_TokenizedCustom cl100k tokenized version of PubChem10M_SELFIES. text1M<n<10M2 likes32 downloads3y agoHugging Face19distil-whisper /whisper_transcriptions_token_idstext100K<n<1M0 likes30 downloads3y agoHugging Face20AmanPriyanshu /Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens Dynamic Topic Modeling Dataset: RedPajama-1T SubSample (100k samples, 1k tokens) 📝Check out the Blog Post This dataset represents a curated subset of the RedPajama-1T Sample dataset, specifically processed for dynamic topic modeling applications. It contains 100,000 samples from the original dataset, with each document limited to the first 1,024 tokens for consistent processing. Dataset Overview Name:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens.textsummarization100K<n<1M8 likes28 downloads2y agoHugging Face21token-opt-org /Token_Optimization_Org AI Safety & Bias Evaluation Conversations Dataset Summary This dataset contains simulated multi-turn conversations designed to evaluate AI language model behavior across two safety-critical domains: self-harm response handling and political bias. Each row represents a single evaluation scenario where an AI model's responses are assessed for safety compliance or neutrality. The dataset is intended to support research and development of safer, less biased AI systems. All… See the full description on the dataset page: https://huggingface.co/datasets/token-opt-org/Token_Optimization_Org.texttext-classificationn<1K0 likes28 downloads5mo agoHugging Face22sajjadanwar0 /token-budgets-catalog Token Budgets — Empirical catalogue and inter-rater reliability data Data for: Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun Incidents, with an Affine-Typed Rust Mitigation as a Case Study. Sajjad Khan, 2026. arXiv:2606.04056 — preprint. This dataset bundles three artefacts referenced in the paper: catalogue — the harvested catalogue of LLM-agent budget-overrun incidents across 21 orchestration frameworks (2023–2026), 167 rows total. The IRR-included… See the full description on the dataset page: https://huggingface.co/datasets/sajjadanwar0/token-budgets-catalog.texttext-classificationn<1K0 likes27 downloads4mo agoHugging Face23blstweb0901 /tokyo-vpn-monitor language: ja en license: mit multilinguality: multilingual size_categories: 1K<n<10K source_datasets: original task_categories: other task_ids: [] pretty_name: Tokyo VPN Speed Monitor Dataset tags: vpn network-monitoring performance-measurement time-series networking internet-measurement automated-testing zero-cost-infrastructure google-apps-script Tokyo VPN Speed Monitor Dataset Dataset Summary The Tokyo VPN Speed Monitor Dataset contains continuous automated… See the full description on the dataset page: https://huggingface.co/datasets/blstweb0901/tokyo-vpn-monitor.tabulartabular-classification1K<n<10K0 likes24 downloads9mo agoHugging Face24hadeelbkh /tokenized-IELTS-writing-task-2-evaluation-DialoGPT-mediumtexttext-generation1K<n<10K2 likes23 downloads1y agoHugging Face25Circularmachines /Batch_indexing_machine_tokenstabular1M<n<10M0 likes22 downloads3y agoHugging Face26NhungNguyen /rams-no-special-tokenstext1K<n<10K0 likes21 downloads1y agoHugging Face27TokenBender /sentence_retrieval_hindi_SFTtext10K<n<100K2 likes18 downloads3y agoHugging Face28Mr-Bridge /paris-vs-tokyo-hotels-2026 Paris vs Tokyo Hotels 2026: Stars & Guest Ratings Hotels in Paris and Tokyo (3 to 5 star), each with hotel name, platform star level and a blended guest rating (Booking, Priceline, Agoda, HotelsCombined). Collected 2026-06-22 by MrBridge. Files kayak_paris_tokyo_2026.csv — 197 hotels (Paris 98, Tokyo 99), the primary blended-OTA source. priceline_paris_tokyo_2026.csv — 62 hotels, an independent Priceline pull used as a robustness check. Columns: city, name… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/paris-vs-tokyo-hotels-2026.tabulartabular-classificationn<1K0 likes18 downloads3mo agoHugging Face29CentificAIResearch /token-optimization AI Safety & Bias Evaluation Conversations Dataset Summary This dataset contains simulated multi-turn conversations designed to evaluate AI language model behavior across two safety-critical domains: self-harm response handling and political bias. Each row represents a single evaluation scenario where an AI model's responses are assessed for safety compliance or neutrality. The dataset is intended to support research and development of safer, less biased AI… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/token-optimization.texttext-classificationn<1K1 likes17 downloads4mo agoHugging Face30Gugu8 /Token-Efficiency token_efficiency_corpus A 2.5 GB CSV corpus teaching LLMs to minimize token usage in their outputs. Progresses from basic filler removal to expert-level nested reasoning compression. Contents verbose_output - The padded, wasteful version of the text efficient_output - The compressed, token-efficient equivalent technique - Compression strategy used subcategory - Specific variant of the technique difficulty - Tier 1 (easiest)… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Token-Efficiency.tabular1M<n<10M0 likes17 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.