CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01acul3 /mc4_und_idfiltered,deduplication MC4-ID from MC4 part undfined text1M<n<10M0 likes484 downloads2y agoHugging Face02Arrrlex /models-under-pressure Models Under Pressure This dataset accompanies the paper Detecting High-Stakes Interactions with Activation Probes, presented at the ICML 2025 Workshop on Actionable Interpretability, accepted to NeurIPS 2025. Overview Every sample is a user-facing LLM interaction labelled as high-stakes or low-stakes. The label reflects whether the conversation involves potentially consequential outcomes (medical advice, legal matters, financial decisions, etc.) vs. routine queries. The… See the full description on the dataset page: https://huggingface.co/datasets/Arrrlex/models-under-pressure.tabulartext-classification10K<n<100K0 likes394 downloads8mo agoHugging Face03theblackcat102 /anime-understanding-dataset Anime Understanding Benchmark (WIP) Evaluate anime knowledge found in existing LLMs. We hope to provide an easy to run evaluation on knowledge understanding in anime/manga. Better understanding in anime/manga knowledge should resulted in task such as waifu role play. Any suggestion is open in discussion tab. Currently in the works [] Eval on popular models such as gpt, hermes, dolphin, llama base model [] Add more metadata regarding of anime/manga year span [] Suggestions… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/anime-understanding-dataset.textquestion-answering1K<n<10K3 likes280 downloads3y agoHugging Face04Undi95 /ConversationChronicles-sharegpt-SHARDEDThis is a sharded version of the PocketDoc/ConversationChronicles-sharegpt dataset, a sharegpt conversion of the jihyoung/ConversationChronicles dataset. All dialogue got fixed (space, coma) and spread across the different relationship available : Relationship Count Ratio Classmates 66,090 33.05% Neighbors 49,521 24.76% Co-workers 28,856 14.43% Mentee and Mentor 16,035 8.02% Husband and Wife 13,486 6.74% Patient and Doctor 6,980 3.49% Parent and Child6,514 3.26%… See the full description on the dataset page: https://huggingface.co/datasets/Undi95/ConversationChronicles-sharegpt-SHARDED.text100K<n<1M11 likes247 downloads3y agoHugging Face05Undi95 /gsm8k-R1text1K<n<10K0 likes232 downloads2y agoHugging Face06KaraKaraWitch /Kurdish-Underwater-Basketweaving-Forum KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum Yes this is a 4chan dataset. THIS CONTAINS TOXIC SHIT like (/POL/) content. YOU HAVE BEEN WARNED. KaraKaraWitch & their company dissolves all responsbilities when using this dataset. Text Sample Note: namedconversation is a modification of OAI's conversation format. While identical, namedconversation is not required to stick to system,user,model/assistant verbs. This allows for a much more varied use… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum.text100K<n<1M5 likes175 downloads2y agoHugging Face07Gporrt /sentiment-understanding-corpusSentiment Corpus Distilling Fine-grained Sentiment Understanding from Large Language Models Fine-grained sentiment analysis (FSA) aims to extract and summarize user opinions from vast opinionated text. Recent studies demonstrate that large language models (LLMs) possess exceptional sentiment understanding capabilities. However, directly deploying LLMs for FSA applications incurs high inference costs. Therefore, this paper investigates the distillation of fine-grained sentiment… See the full description on the dataset page: https://huggingface.co/datasets/Gporrt/sentiment-understanding-corpus.tabular1M<n<10M1 likes163 downloads1y agoHugging Face08Undi95 /R1-RP-ShareGPT3Entire dataset of Mistral Thinker. V3. text10K<n<100K8 likes154 downloads2y agoHugging Face09undertheseanlp /UTS2017_Bank UTS2017_Bank Dataset Dataset Description Dataset Summary The UTS2017_Bank dataset is a comprehensive Vietnamese banking domain dataset containing customer feedback and reviews about banking services. It contains 2,471 annotated examples (1,977 train, 494 test) with both aspect labels and sentiment annotations. The dataset supports multiple NLP tasks including aspect classification, sentiment analysis, and aspect-based sentiment analysis in the Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS2017_Bank.texttext-classification1K<n<10K2 likes89 downloads1y agoHugging Face10WeMake /Intelligent-Content-Understanding Intelligent Content Understanding Empowering Advanced Thinking, Deep Understanding, Diverse Perspectives, and Creative Solutions Across Disciplines By fostering a richly interconnected knowledge ecosystem, ICU (Intelligent Content Understanding) aims to elevate language models to unparalleled heights of understanding, reasoning, and innovation. This ambitious project lays the groundwork for developing an 'internal knowledge map' within language models, enabling… See the full description on the dataset page: https://huggingface.co/datasets/WeMake/Intelligent-Content-Understanding.texttext-generation1K<n<10K6 likes76 downloads1y agoHugging Face11undertheseanlp /UVB-v0.1 UVB - Underthesea Vietnamese Books Dataset A collection of 447 Vietnamese books with full text content and Goodreads metadata for NLP research. Dataset Summary UVB (Underthesea Vietnamese Books) is a dataset containing 447 Vietnamese books with full text content, mapped to Goodreads for metadata enrichment including genres, ratings, and publication years. The dataset is designed for Vietnamese language model training, text generation, and other NLP tasks.… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVB-v0.1.tabulartext-generationn<1K0 likes65 downloads8mo agoHugging Face12Undi95 /Filtered_Tree_Of_Thoughts_BASE_24kOriginal dataset : https://huggingface.co/datasets/terrycraddock/Tree_Of_Thoughts_BASE_24k <output> and </output> filtered Added 2 new special token for llama 3.1 & llama 3.3 : <|start_thinking|> and <|end_thinking|> Filtered reply where the thinking never ended or never started. Usage in axolotl : datasets: - path: Undi95/Filtered_Tree_Of_Thoughts_BASE_24k type: alpaca_chat.load_qa conversation: llama3 Prompt template usage :… See the full description on the dataset page: https://huggingface.co/datasets/Undi95/Filtered_Tree_Of_Thoughts_BASE_24k.text10K<n<100K9 likes64 downloads2y agoHugging Face13Undi95 /andrijdavid_roleplay-conversation-sharegptShareGPT formated dataset from andrijdavid/roleplay-conversation It miss some data at the end because I was too lazy to continue, there was too much things to modify, I do it by hand/notepad++/regex lmao text10K<n<100K13 likes51 downloads3y agoHugging Face14lucasjin /undefineddtext100K<n<1M0 likes48 downloads2y agoHugging Face15mstyslavity /philosophy_undergradtexttext-generation100K<n<1M0 likes48 downloads7mo agoHugging Face16XuehangCang /china-undergraduate-majors-2026 普通高等学校本科专业目录(2026年) 本数据集收录了中华人民共和国教育部于 2026 年 4 月发布的《普通高等学校本科专业目录》,以结构化 JSON 格式提供 数据概览 项目 数量 学科门类 13 专业类 93 专业总数 875 特设专业(代码后加"T") 523 国家控制布点专业(代码后加"K") 171 13 个学科门类 代码 学科门类 01 哲学 02 经济学 03 法学 04 教育学 05 文学 06 历史学 07 理学 08 工学 09 农学 10 医学 12 管理学 13 艺术学 14 交叉学科 关于本目录 《普通高等学校本科专业目录》是高等教育工作的基本指导性文件之一,规定专业划分、名称及所属门类,是设置和调整专业、实施人才培养、安排招生、授予学位、指导就业,进行教育统计和人才需求预测等工作的重要依据,专业目录每年更新发布… See the full description on the dataset page: https://huggingface.co/datasets/XuehangCang/china-undergraduate-majors-2026.textn<1K0 likes47 downloads5mo agoHugging Face17Undi95 /toxic-dpo-v0.1-NoWarningtextn<1K25 likes44 downloads3y agoHugging Face18reinaldog /repro-understanding-sam-through-minimax-perspective-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes42 downloads2mo agoHugging Face19PsychiatryAgentBench25 /Health_Information_Seeking_under_Limited_Evidence Health Information Seeking under Limited Evidence (HISLE) HISLE is a clinically informed benchmark for evaluating LLM-based agents responding to incomplete mental-health information needs. File Records Contents matched_pairs_47.jsonl 47 Matched Chinese–English scenario pairs matched_variants_3290.jsonl 3,290 Query variants for the matched scenarios coverage_originals_24.jsonl 24 Coverage-expansion queries coverage_variants_840.jsonl 840 Query variants for… See the full description on the dataset page: https://huggingface.co/datasets/PsychiatryAgentBench25/Health_Information_Seeking_under_Limited_Evidence.tabular1K<n<10K0 likes40 downloads17d agoHugging Face20eli-underwood /rings-multiobj101textn<1K0 likes40 downloads18d agoHugging Face21tomyimkc /repro-dimension-independent-convergence-of-underdamped-langevin-monte-carlo-in-kl-dive-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes38 downloads2mo agoHugging Face22randomyellowdude /repro-understanding-lora-as-knowledge-memory-an-empirical-analysis-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes36 downloads2mo agoHugging Face23eli-underwood /rings-chair101textn<1K0 likes36 downloads18d agoHugging Face24eli-underwood /rings-primitive101textn<1K0 likes36 downloads18d agoHugging Face25underscore2 /bluesky_tpottext1K<n<10K1 likes33 downloads1y agoHugging Face26Undi95 /toxic-dpo-v0.1-sharegptUPDATE: Merged the NoWarning into a real DPO for later use. Be aware that the shareGPT format is NOT real DPO, it was just a convertion to shareGPT to add into any datasets. If you want to do a REAL DPO train, use this file: toxic-dpo-NoWarning.json. DISCLAIMER : I'M NOT THE AUTHOR OF THIS DATASET. ALL CREDIT GO TO unalignment repo. ORIGINAL DATASET: unalignment/toxic-dpo-v0.1 I just converted/modified the dataset! Only the accepted replies was taken for the shareGPT format!… See the full description on the dataset page: https://huggingface.co/datasets/Undi95/toxic-dpo-v0.1-sharegpt.textn<1K18 likes32 downloads3y agoHugging Face27eli-underwood /plato-rings-6ktext1K<n<10K0 likes32 downloads12d agoHugging Face28sabaridsnfuji /repro-box-thirding-anytime-best-arm-identification-under-insufficient-sampling Box Thirding (B3): Anytime Best Arm Identification under Insufficient Sampling Reproduction of ICML 2026 paper (OpenReview: XoONWh8fbL) Tags trackio trackio-logbook open-experiment icml2026-repro paper-XoONWh8fbL textn<1K0 likes31 downloads2mo agoHugging Face29Undi95 /Weyaxi-humanish-dpo-project-noemojitext1K<n<10K18 likes27 downloads2y agoHugging Face30eli-underwood /plato-multiobjtext1K<n<10K0 likes26 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.