datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
models-under-pressure
Models Under Pressure
This dataset accompanies the paper Detecting High-Stakes Interactions with Activation Probes, presented at the ICML 2025 Workshop on Actionable Interpretability, accepted to NeurIPS 2025.
Overview
Every sample is a user-facing LLM interaction labelled as high-stakes or low-stakes. The label reflects whether the conversation involves potentially consequential outcomes (medical advice, legal matters, financial decisions, etc.) vs. routine queries.
The… See the full description on the dataset page: https://huggingface.co/datasets/Arrrlex/models-under-pressure.sentiment-understanding-corpusSentiment Corpus
Distilling Fine-grained Sentiment Understanding from Large Language Models
Fine-grained sentiment analysis (FSA) aims to extract and summarize user opinions from vast opinionated text. Recent studies demonstrate that large language models (LLMs) possess exceptional sentiment understanding capabilities. However, directly deploying LLMs for FSA applications incurs high inference costs. Therefore, this paper investigates the distillation of fine-grained sentiment… See the full description on the dataset page: https://huggingface.co/datasets/Gporrt/sentiment-understanding-corpus.UVB-v0.1
UVB - Underthesea Vietnamese Books Dataset
A collection of 447 Vietnamese books with full text content and Goodreads metadata for NLP research.
Dataset Summary
UVB (Underthesea Vietnamese Books) is a dataset containing 447 Vietnamese books with full text content, mapped to Goodreads for metadata enrichment including genres, ratings, and publication years. The dataset is designed for Vietnamese language model training, text generation, and other NLP tasks.… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVB-v0.1.repro-understanding-sam-through-minimax-perspective-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Health_Information_Seeking_under_Limited_Evidence
Health Information Seeking under Limited Evidence (HISLE)
HISLE is a clinically informed benchmark for evaluating LLM-based agents responding to incomplete mental-health information needs.
File
Records
Contents
matched_pairs_47.jsonl
47
Matched Chinese–English scenario pairs
matched_variants_3290.jsonl
3,290
Query variants for the matched scenarios
coverage_originals_24.jsonl
24
Coverage-expansion queries
coverage_variants_840.jsonl
840
Query variants for… See the full description on the dataset page: https://huggingface.co/datasets/PsychiatryAgentBench25/Health_Information_Seeking_under_Limited_Evidence.repro-dimension-independent-convergence-of-underdamped-langevin-monte-carlo-in-kl-dive-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-understanding-lora-as-knowledge-memory-an-empirical-analysis-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Undi95__MG-FinalMix-72B-details
Dataset Card for Evaluation run of Undi95/MG-FinalMix-72B
Dataset automatically created during the evaluation run of model Undi95/MG-FinalMix-72B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Undi95__MG-FinalMix-72B-details.TheHierophant__Underground-Cognitive-V0.3-test-details
Dataset Card for Evaluation run of TheHierophant/Underground-Cognitive-V0.3-test
Dataset automatically created during the evaluation run of model TheHierophant/Underground-Cognitive-V0.3-test
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TheHierophant__Underground-Cognitive-V0.3-test-details.Undi95__Phi4-abliterated-details
Dataset Card for Evaluation run of Undi95/Phi4-abliterated
Dataset automatically created during the evaluation run of model Undi95/Phi4-abliterated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Undi95__Phi4-abliterated-details.
