li-lab/Omanic
Omanic The dataset contains two splits: OmanicSynth: 10,296 machine-generated training examples. OmanicBench: 967 expert-reviewed, human-annotated evaluation examples. Each row is a single 4-hop reasoning instance. For more details, please refer to the paper: Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models. Data Fields Each example contains the following top-level keys: id: A unique example identifier. single_hop: A list of… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/Omanic.
0113
