datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mc4_und_idfiltered,deduplication MC4-ID from MC4 part undfined
models-under-pressure
Models Under Pressure
This dataset accompanies the paper Detecting High-Stakes Interactions with Activation Probes, presented at the ICML 2025 Workshop on Actionable Interpretability, accepted to NeurIPS 2025.
Overview
Every sample is a user-facing LLM interaction labelled as high-stakes or low-stakes. The label reflects whether the conversation involves potentially consequential outcomes (medical advice, legal matters, financial decisions, etc.) vs. routine queries.
The… See the full description on the dataset page: https://huggingface.co/datasets/Arrrlex/models-under-pressure.anime-understanding-dataset
Anime Understanding Benchmark (WIP)
Evaluate anime knowledge found in existing LLMs. We hope to provide an easy to run evaluation on knowledge understanding in anime/manga. Better understanding in anime/manga knowledge should resulted in task such as waifu role play.
Any suggestion is open in discussion tab.
Currently in the works
[] Eval on popular models such as gpt, hermes, dolphin, llama base model
[] Add more metadata regarding of anime/manga year span
[] Suggestions… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/anime-understanding-dataset.ConversationChronicles-sharegpt-SHARDEDThis is a sharded version of the PocketDoc/ConversationChronicles-sharegpt dataset, a sharegpt conversion of the jihyoung/ConversationChronicles dataset.
All dialogue got fixed (space, coma) and spread across the different relationship available :
Relationship
Count
Ratio
Classmates
66,090
33.05%
Neighbors
49,521
24.76%
Co-workers
28,856
14.43%
Mentee and Mentor
16,035
8.02%
Husband and Wife
13,486
6.74%
Patient and Doctor
6,980
3.49%
Parent and Child6,514
3.26%… See the full description on the dataset page: https://huggingface.co/datasets/Undi95/ConversationChronicles-sharegpt-SHARDED.gsm8k-R1Kurdish-Underwater-Basketweaving-Forum
KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum
Yes this is a 4chan dataset. THIS CONTAINS TOXIC SHIT like (/POL/) content. YOU HAVE BEEN WARNED.
KaraKaraWitch & their company dissolves all responsbilities when using this dataset.
Text Sample
Note: namedconversation is a modification of OAI's conversation format. While identical, namedconversation is not required to stick to system,user,model/assistant verbs. This allows for a much more varied use… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum.sentiment-understanding-corpusSentiment Corpus
Distilling Fine-grained Sentiment Understanding from Large Language Models
Fine-grained sentiment analysis (FSA) aims to extract and summarize user opinions from vast opinionated text. Recent studies demonstrate that large language models (LLMs) possess exceptional sentiment understanding capabilities. However, directly deploying LLMs for FSA applications incurs high inference costs. Therefore, this paper investigates the distillation of fine-grained sentiment… See the full description on the dataset page: https://huggingface.co/datasets/Gporrt/sentiment-understanding-corpus.R1-RP-ShareGPT3Entire dataset of Mistral Thinker.
V3.
UTS2017_Bank
UTS2017_Bank Dataset
Dataset Description
Dataset Summary
The UTS2017_Bank dataset is a comprehensive Vietnamese banking domain dataset containing customer feedback and reviews about banking services. It contains 2,471 annotated examples (1,977 train, 494 test) with both aspect labels and sentiment annotations. The dataset supports multiple NLP tasks including aspect classification, sentiment analysis, and aspect-based sentiment analysis in the Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS2017_Bank.Intelligent-Content-Understanding
Intelligent Content Understanding
Empowering Advanced Thinking, Deep Understanding, Diverse Perspectives, and Creative Solutions Across Disciplines
By fostering a richly interconnected knowledge ecosystem, ICU (Intelligent Content Understanding) aims to elevate language models to unparalleled heights of understanding, reasoning, and innovation.
This ambitious project lays the groundwork for developing an 'internal knowledge map' within language models, enabling… See the full description on the dataset page: https://huggingface.co/datasets/WeMake/Intelligent-Content-Understanding.UVB-v0.1
UVB - Underthesea Vietnamese Books Dataset
A collection of 447 Vietnamese books with full text content and Goodreads metadata for NLP research.
Dataset Summary
UVB (Underthesea Vietnamese Books) is a dataset containing 447 Vietnamese books with full text content, mapped to Goodreads for metadata enrichment including genres, ratings, and publication years. The dataset is designed for Vietnamese language model training, text generation, and other NLP tasks.… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVB-v0.1.Filtered_Tree_Of_Thoughts_BASE_24kOriginal dataset : https://huggingface.co/datasets/terrycraddock/Tree_Of_Thoughts_BASE_24k
<output> and </output> filtered
Added 2 new special token for llama 3.1 & llama 3.3 : <|start_thinking|> and <|end_thinking|>
Filtered reply where the thinking never ended or never started.
Usage in axolotl :
datasets:
- path: Undi95/Filtered_Tree_Of_Thoughts_BASE_24k
type: alpaca_chat.load_qa
conversation: llama3
Prompt template usage :… See the full description on the dataset page: https://huggingface.co/datasets/Undi95/Filtered_Tree_Of_Thoughts_BASE_24k.andrijdavid_roleplay-conversation-sharegptShareGPT formated dataset from andrijdavid/roleplay-conversation
It miss some data at the end because I was too lazy to continue, there was too much things to modify, I do it by hand/notepad++/regex lmao
undefineddphilosophy_undergradchina-undergraduate-majors-2026
普通高等学校本科专业目录(2026年)
本数据集收录了中华人民共和国教育部于 2026 年 4 月发布的《普通高等学校本科专业目录》,以结构化 JSON 格式提供
数据概览
项目
数量
学科门类
13
专业类
93
专业总数
875
特设专业(代码后加"T")
523
国家控制布点专业(代码后加"K")
171
13 个学科门类
代码
学科门类
01
哲学
02
经济学
03
法学
04
教育学
05
文学
06
历史学
07
理学
08
工学
09
农学
10
医学
12
管理学
13
艺术学
14
交叉学科
关于本目录
《普通高等学校本科专业目录》是高等教育工作的基本指导性文件之一,规定专业划分、名称及所属门类,是设置和调整专业、实施人才培养、安排招生、授予学位、指导就业,进行教育统计和人才需求预测等工作的重要依据,专业目录每年更新发布… See the full description on the dataset page: https://huggingface.co/datasets/XuehangCang/china-undergraduate-majors-2026.toxic-dpo-v0.1-NoWarningrepro-understanding-sam-through-minimax-perspective-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Health_Information_Seeking_under_Limited_Evidence
Health Information Seeking under Limited Evidence (HISLE)
HISLE is a clinically informed benchmark for evaluating LLM-based agents responding to incomplete mental-health information needs.
File
Records
Contents
matched_pairs_47.jsonl
47
Matched Chinese–English scenario pairs
matched_variants_3290.jsonl
3,290
Query variants for the matched scenarios
coverage_originals_24.jsonl
24
Coverage-expansion queries
coverage_variants_840.jsonl
840
Query variants for… See the full description on the dataset page: https://huggingface.co/datasets/PsychiatryAgentBench25/Health_Information_Seeking_under_Limited_Evidence.rings-multiobj101repro-dimension-independent-convergence-of-underdamped-langevin-monte-carlo-in-kl-dive-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-understanding-lora-as-knowledge-memory-an-empirical-analysis-traces
Agent traces
Agent sessions published from a Trackio Logbook.
rings-chair101rings-primitive101bluesky_tpottoxic-dpo-v0.1-sharegptUPDATE: Merged the NoWarning into a real DPO for later use. Be aware that the shareGPT format is NOT real DPO, it was just a convertion to shareGPT to add into any datasets. If you want to do a REAL DPO train, use this file: toxic-dpo-NoWarning.json.
DISCLAIMER : I'M NOT THE AUTHOR OF THIS DATASET.
ALL CREDIT GO TO unalignment repo.
ORIGINAL DATASET: unalignment/toxic-dpo-v0.1
I just converted/modified the dataset! Only the accepted replies was taken for the shareGPT format!… See the full description on the dataset page: https://huggingface.co/datasets/Undi95/toxic-dpo-v0.1-sharegpt.plato-rings-6krepro-box-thirding-anytime-best-arm-identification-under-insufficient-sampling
Box Thirding (B3): Anytime Best Arm Identification under Insufficient Sampling
Reproduction of ICML 2026 paper (OpenReview: XoONWh8fbL)
Tags
trackio
trackio-logbook
open-experiment
icml2026-repro
paper-XoONWh8fbL
Weyaxi-humanish-dpo-project-noemojiplato-multiobj
