datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SMART
SMART: Evaluating LLMs’ Mathematical Reasoning via a Human Cognitive Process-Inspired Benchmark
SMART is a fine-grained benchmark for evaluating large language models (LLMs) on mathematical reasoning from a human cognitive process perspective. Instead of evaluating only the final answer, SMART decomposes mathematical problem solving into four cognitive dimensions inspired by Pólya’s problem-solving theory:
Semantic Understanding
Mathematical Reasoning
Arithmetic Computation… See the full description on the dataset page: https://huggingface.co/datasets/ewdfd/SMART.smartroute-rag-synthetic-routing-benchmark-5000
🧭 SmartRoute-RAG Synthetic Routing Benchmark 5000
A publication-scale benchmark for evaluating when to retrieve — not just what to answer.
5,000 stratified questions · 10 benchmark-style subsets · 13 question types · binary routing labelsBuilt for the SmartRoute-RAG research line: false-skip-aware, safety-constrained adaptive retrieval.
🎯 Why this dataset exists
Most RAG benchmarks measure answer quality after retrieval. They rarely tell you whether the… See the full description on the dataset page: https://huggingface.co/datasets/pr0mila-gh0sh/smartroute-rag-synthetic-routing-benchmark-5000.SmartHome-Device-QAqa-dataset-20250127arc-trThis Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish benchmarks to evaluate the performance of LLM's Produced in the Turkish Language.
Dataset Card for arc-tr
malhajar/arc-tr is a translated version of arc aimed specifically to be used in the OpenLLMTurkishLeaderboard
This Dataset contains rigid tests extracted from the paper Think you have Solved Question Answering?
Developed by: Mohamad Alhajar
Data… See the full description on the dataset page: https://huggingface.co/datasets/smart011/arc-tr.
