datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPQA-trajectory
GPQA Trajectory Dataset
Model-generated solution trajectories for GPQA (Graduate-Level Google-Proof Q&A), a multiple-choice benchmark of graduate-level physics, chemistry, and biology questions. Each row is one model response to a single problem, including the hidden chain-of-thought (when available) and the final response.
Dataset Summary
Split
Rows
Unique Problems
Model(s)
Has reasoning_content
Accuracy
train (non-diamond)
410
334
deepseek-r1
Yes
100%… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/GPQA-trajectory.less-is-moe-gpqa-diamond-evaluation
Less-is-MoE GPQA-Diamond evaluation set
This private dataset stores the 198-question GPQA-Diamond evaluation file used
by the MoE-Honing evaluation format.
Upstream source: Idavidrein/gpqa, config gpqa_diamond
Upstream revision: 633f5ee89ab8ad4522a9f850766b73f62147ffdd
Split: test
Rows: 198
SHA-256: d5b0d6dad6c1993a8cb17fd7aa635fbc5e8b684ae42b5b6e80d15467b399eec3
Fields: problem, solution, domain
The problem field contains the formatted four-choice prompt, solution stores
the… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-gpqa-diamond-evaluation.
