Deduction
proofwriter-deduction-balancedA processed subset of the OWA section of the ProofWriter dataset.
Each train/test split contains 300 entries, each of which has a unique set of theories and a single question for those theories.
Both splits are balanced so that the depth of the proof required to answer the question varies evenly between 0-5 (50 entries each), and the labels are balanced (100 each).
'Unknown' labels have been replaced by 'Uncertain' to match other datasets.
faithfulness-logical_deduction_five_objectsentity-deduction-arena
Entity-Deduction Arena (EDA)
This dataset complements the paper Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games, presented in ACL 2024 main conference.
The main repo can be found at https://github.com/apple/ml-entity-deduction-arena
Motivation
There is a demand to assessing the capability of LLM to clarify with questions in order to effectively resolve ambiguities, when confronted with vague queries.
This capability demands a sophisticated… See the full description on the dataset page: https://huggingface.co/datasets/yizheapple/entity-deduction-arena.fineweb-edu-json-schema-deduction
FineWeb-Edu json-schema-deduction
Source: fineweb-edu dataset.
Task: JSON schema deduction.
5,000 entries from fineweb-edu dataset
btw, every single key in the schema is unique. The model reasoning was high. The ontology went too deep haha.
It generated over 51,000 unique keys across 5,000 documents. it basically baked raw text directly into the structural keys.
however! json is 100% valid and correct so theres that
Columns are raw_text and schema
bbh-logical-deduction-seven-objects-pllogical-deduction-filtered-qwen3-0.6b-no-think-2048
