datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
debate-tracking-v3
Debate Tracking Dataset v3
Training data from 30 competitive debates (10 topics × 3 judges) with multi-response scoring.
Dataset Description
Each row represents a single LLM call during debate generation, with multiple response variations scored by Claude Sonnet.
Statistics
Debates: 30
Topics: 10 diverse IPDA debate resolutions
Judges: 3 different judge profiles (lay, parent, coach)
Training Examples: 1,816 calls
Winner Distribution: AFF 33%, NEG 67%… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-tracking-v3.hanabi-fireworks-state-tracking
FIREWORKS: Hanabi belief-state reconstruction
Strict hidden-belief reconstruction for Hanabi. Each example gives a previous belief state
plus the actions taken since, and asks for the updated per-card possibility sets.
10,185 examples (9,780 unique prompts)
2-5 player games, seeds 101-110 (evaluation seeds 1001-1010 are held out)
Fields: id, meta (num_players, seed, turn, observer, log), prompt, target
Assembled from two labeling passes (GPT-4.1-mini: 7,232 rows; Grok-3-mini: 2… See the full description on the dataset page: https://huggingface.co/datasets/Mahesh111000/hanabi-fireworks-state-tracking.textworld-state-tracking
TextWorld State Tracking Dataset
Dataset Description
This dataset contains 27,145 examples for evaluating language models' ability to track object states across narrative contexts. The data is derived from TextWorld, an interactive text-based game environment, and formatted for state tracking evaluation.
Dataset Summary
Each example presents a narrative context (a sequence of game observations and actions) along with:
An entity (object) being tracked
Its… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/textworld-state-tracking.assumption-tracking-dependency-awareness-v01
"awareness.csv"
Cardinal Meta Dataset 2Assumption Tracking and Dependency Awareness
Purpose
Test whether the model names assumptions
Test whether conclusions track their dependencies
Test whether removing an assumption collapses the claim
Core question
What must be true for this to be true
Why this is meta
The dataset does not test domain facts
It tests whether the model keeps structure attached to claims
It sits above domains because every domain rests on assumptions
What it… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/assumption-tracking-dependency-awareness-v01.
