datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
proofwriter-deduction-balancedA processed subset of the OWA section of the ProofWriter dataset.
Each train/test split contains 300 entries, each of which has a unique set of theories and a single question for those theories.
Both splits are balanced so that the depth of the proof required to answer the question varies evenly between 0-5 (50 entries each), and the labels are balanced (100 each).
'Unknown' labels have been replaced by 'Uncertain' to match other datasets.
faithfulness-logical_deduction_five_objectsentity-deduction-arena
Entity-Deduction Arena (EDA)
This dataset complements the paper Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games, presented in ACL 2024 main conference.
The main repo can be found at https://github.com/apple/ml-entity-deduction-arena
Motivation
There is a demand to assessing the capability of LLM to clarify with questions in order to effectively resolve ambiguities, when confronted with vague queries.
This capability demands a sophisticated… See the full description on the dataset page: https://huggingface.co/datasets/yizheapple/entity-deduction-arena.fineweb-edu-json-schema-deduction
FineWeb-Edu json-schema-deduction
Source: fineweb-edu dataset.
Task: JSON schema deduction.
5,000 entries from fineweb-edu dataset
btw, every single key in the schema is unique. The model reasoning was high. The ontology went too deep haha.
It generated over 51,000 unique keys across 5,000 documents. it basically baked raw text directly into the structural keys.
however! json is 100% valid and correct so theres that
Columns are raw_text and schema
bbh-logical-deduction-seven-objects-pllogical-deduction-filtered-qwen3-0.6b-no-think-2048Entity-deduction-arena-20-questions
twentyquestions
The twentyquestios dataset provides
20 Questions style questions
and answers collected from real games of twentyquestions between people on
Mechanical Turk.
In this directory, you'll find the following files:
twentyquestions-train.jsonl
twentyquestions-dev.jsonl
twentyquestions-test.jsonl
twentyquestions-all.jsonl
The files represent the train, dev, and test splits as well as all the data
together before being split. Train, dev, and test were split by dividing up… See the full description on the dataset page: https://huggingface.co/datasets/jtv199/Entity-deduction-arena-20-questions.bayesian-social-deduction
Bayesian Social Deduction Dataset
Project Page | Arxiv | Github
Dataset Description
This dataset contains a collection of game logs from Avalon social deduction games, generated for the "Bayesian Social Deduction with Graph-Informed Language Models" paper. The dataset includes games played by various agents, including humans, and different AI models, providing a rich resource for analyzing strategic communication, deception, and cooperation.
The dataset is organized into… See the full description on the dataset page: https://huggingface.co/datasets/shahabrahimirad/bayesian-social-deduction.logical-deduction-filtered-qwen3-0.6b-2048faithfulness-logical_deduction_five_objects-arbbridges_5x5de_deduction_v3proofwriter-deduction-balancedA processed subset of the OWA section of the ProofWriter dataset.
Each train/test split contains 300 entries, each of which has a unique set of theories and a single question for those theories.
Both splits are balanced so that the depth of the proof required to answer the question varies evenly between 0-5 (50 entries each), and the labels are balanced (100 each).
'Unknown' labels have been replaced by 'Uncertain' to match other datasets.
test_helpsteer2_prefrence_sfr_llama31_repsonse_deductionbbh-logical-deduction-seven-objects-pl-100From_Extraction_to_Deduction
