datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPQA-diamond-ClaudeR1
Dataset Card for GPQA Diamond Reasoning Benchmark
Dataset Details
Dataset Description
A benchmark dataset for evaluating hybrid AI architectures, comparing reasoning-augmented LLMs (DeepSeek R1) against standalone models (Claude Sonnet 3.5). Contains 198 physics questions with:
Ground truth answers and explanations
Model responses from multiple architectures
Granular token usage and cost metrics
Difficulty metadata and domain categorization
Curated by: LLM… See the full description on the dataset page: https://huggingface.co/datasets/spawn99/GPQA-diamond-ClaudeR1.my-awesome-dataset
