CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ewdfd /SMART SMART: Evaluating LLMs’ Mathematical Reasoning via a Human Cognitive Process-Inspired Benchmark SMART is a fine-grained benchmark for evaluating large language models (LLMs) on mathematical reasoning from a human cognitive process perspective. Instead of evaluating only the final answer, SMART decomposes mathematical problem solving into four cognitive dimensions inspired by Pólya’s problem-solving theory: Semantic Understanding Mathematical Reasoning Arithmetic Computation… See the full description on the dataset page: https://huggingface.co/datasets/ewdfd/SMART.tabularquestion-answering1K<n<10K2 likes34 downloads5mo agoHugging Face02pr0mila-gh0sh /smartroute-rag-synthetic-routing-benchmark-5000 🧭 SmartRoute-RAG Synthetic Routing Benchmark 5000 A publication-scale benchmark for evaluating when to retrieve — not just what to answer. 5,000 stratified questions · 10 benchmark-style subsets · 13 question types · binary routing labelsBuilt for the SmartRoute-RAG research line: false-skip-aware, safety-constrained adaptive retrieval. 🎯 Why this dataset exists Most RAG benchmarks measure answer quality after retrieval. They rarely tell you whether the… See the full description on the dataset page: https://huggingface.co/datasets/pr0mila-gh0sh/smartroute-rag-synthetic-routing-benchmark-5000.textquestion-answering1K<n<10K0 likes19 downloads2mo agoHugging Face03USER9724 /SmartHome-Device-QAtextquestion-answering1K<n<10K1 likes9 downloads2y agoHugging Face04smartyalgo /qa-dataset-20250127textquestion-answeringn<1K0 likes9 downloads2y agoHugging Face05smart011 /arc-trThis Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish benchmarks to evaluate the performance of LLM's Produced in the Turkish Language. Dataset Card for arc-tr malhajar/arc-tr is a translated version of arc aimed specifically to be used in the OpenLLMTurkishLeaderboard This Dataset contains rigid tests extracted from the paper Think you have Solved Question Answering? Developed by: Mohamad Alhajar Data… See the full description on the dataset page: https://huggingface.co/datasets/smart011/arc-tr.textquestion-answering1K<n<10K0 likes8 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.