datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ethical-Reasoning-in-Mental-Health-v1This repository contains the dataset for the paper EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI.
Overview
Ethical-Reasoning-in-Mental-Health-v1 (EthicsMH) is a carefully curated dataset focused on ethical decision-making scenarios in mental health contexts.This dataset captures the complexity of real-world dilemmas faced by therapists, psychiatrists, and AI systems when navigating critical issues such as confidentiality, autonomy, and bias.
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/UVSKKR/Ethical-Reasoning-in-Mental-Health-v1.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.morse-500
MORSE-500 Benchmark
🔥 News
May 15, 2025: We release MORSE-500, 500 programmatically generated videos across six reasoning categories: abstract, mathematical, physical, planning, spatial, and temporal, to stress-test multimodal reasoning. Frontier models including OpenAI o3 and Gemini 2.5 Pro score lower than… See the full description on the dataset page: https://huggingface.co/datasets/video-reasoning/morse-500.sober_reasoning
🧠 Sober Reasoning: Evaluation Logs
This repository hosts evaluation logs and outputs from our paper:
"A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility"
📄 Paper📊 Leaderboard💻 Evaluation Code
🗂️ Repository Structure
Evaluation logs are organized by the cluster used during inference to highlight hardware-induced variance in model performance (see Section 3.3 of the paper).
sober_reasoning/
├── cluster_A/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/bethgelab/sober_reasoning.agentic-reasoning-benchmark
Agentic & Reasoning Benchmark (ARB) – Expanded
Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning.
Überblick
Eigenschaft
Wert
Anzahl Beispiele
2.550
Kategorien
8
Schwierigkeitsgrade
easy / medium / hard
Formate
CSV + JSON
Reproduzierbarkeit
Generator-Skript (seed=42) enthalten
Lizenz
CC-BY-4.0
Kategorien
Kategorie
Anzahl
Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.easy_turkish_math_reasoning
Easy Turkish Math Reasoning
Dataset Summary
The Easy Turkish Math Reasoning dataset is the first phase of a multi-stage curriculum learning pipeline designed to enhance the reasoning abilities of compact language models. This dataset focuses on elementary-level arithmetic and logic problems in Turkish, serving as a warm-up stage for supervised fine-tuning (SFT).
Use Case
Primarily used for:
Bootstrapping reasoning ability in Turkish for compact LLMs.
Phase 1… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/easy_turkish_math_reasoning.medium_turkish_math_reasoning
Dataset Summary
The Medium Turkish Math Reasoning dataset is Phase 2 of a curriculum learning pipeline to teach compact models multi-step reasoning in Turkish. It includes moderately difficult math problems involving multiple reasoning steps, such as two-part arithmetic, comparisons, and logical reasoning.
Use Case
This dataset is ideal for:
Continuing SFT after foundational training with simpler problems.
Bridging the gap between basic arithmetic and complex GSM8K-style… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/medium_turkish_math_reasoning.multilevel-legal-reasoning
Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations
Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi
Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai.
🧭 Purpose and Scope
The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains… See the full description on the dataset page: https://huggingface.co/datasets/lucsaint/Deepseek-V4-Reasoning-Code-2500.logicwaver-reasoning-v1
LogicWaver Adversarial Reasoning Benchmark
Adversarial Semantic Reasoning Benchmark — 20 handcrafted odd-one-out puzzles where SOTA LLMs fail but humans succeed. For LLM evaluation & Chain-of-Thought probing.
🚀 Live Demo: https://huggingface.co/spaces/Eviezonr08/logicwaver-demo
▶️ Try 20 puzzles interactively - no install!
Files
Reasoning_Puzzle_without_proline.csv — 20 puzzles for evaluation
with_proline/Reasoning_Puzzle_proline.csv — Same 20 + pro_line… See the full description on the dataset page: https://huggingface.co/datasets/Eviezonr08/logicwaver-reasoning-v1.symfony-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Symfony
Documentation Data Source Link: https://symfony.com/doc/current/index.html
Data Source License: https://github.com/symfony/symfony?tab=MIT-1-ov-file#readme
Data Source Authors: Symfony SAS
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
Arabic-Optimized-Reasoning-Dataset
Arabic Optimized Reasoning Dataset
Dataset Name: Arabic Optimized ReasoningLicense: Apache-2.0Formats: CSVSize: 1600 rowsBase Dataset: cognitivecomputations/dolphin-r1Libraries Used: Datasets, Dask, Croissant
Overview
The Arabic Optimized Reasoning Dataset helps AI models get better at reasoning in Arabic. While AI models are good at many tasks, they often struggle with reasoning in languages other than English. This dataset helps fix this problem by:
Using fewer tokens… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/Arabic-Optimized-Reasoning-Dataset.langchain-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: langchain
Documentation Data Source Link: https://python.langchain.com/docs/introduction/
Data Source License: https://github.com/langchain-ai/langchain/blob/master/LICENSE
Data Source Authors: Observable AI Benchmarks by Data Agents © 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
tensorflow-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: tensorflow
Data Source Link: https://www.tensorflow.org/api_docs
Data Source License: https://github.com/tensorflow/tensorflow/blob/master/LICENSE
Data Source Authors: Google
AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
gradio-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: gradio
Documentation Data Source Link: https://www.gradio.app/docs
Data Source License: https://github.com/gradio-app/gradio/blob/main/LICENSE
Data Source Authors: Observable AI Benchmarks by Data Agents © 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
scikit-learn-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: scikit-learn
Data Source Link: https://scikit-learn.org/stable/index.html
Data Source License: https://github.com/scikit-learn/scikit-learn/blob/main/COPYING
Data Source Authors: scikit-learn contributors
AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
sino-xenic-reasoning-gap-dataset
Sino-Xenic Reasoning Gap Dataset
A comprehensive evaluation dataset for testing Large Language Models' understanding of Sino-Xenic linguistic phenomena across Chinese, Japanese, Korean, and Vietnamese.
Dataset Overview
Total Samples: 297
Languages: Chinese, Japanese, Korean, Vietnamese
Categories: 11
Task Types: Surface-level and Deep Structural
Categories
Chinese Idioms (27 samples) - Understanding Chinese idioms and their cultural meanings
Chinese… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhasia/sino-xenic-reasoning-gap-dataset.mmcv-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: mmcv
Data Source Link: https://mmcv.readthedocs.io/en/latest/
Data Source License: https://github.com/open-mmlab/mmcv?tab=Apache-2.0-1-ov-file#readme
Data Source Authors: OpenMMLab Computer Vision Foundation
AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
oncology-financial-reasoning-india
🩺 Medzz-AI: Oncology & Financial Reasoning (India)
Status: Active | Context: Indian Healthcare | Focus: Clinical + Economic Logic
👋 The Problem: Why Current Medical AI Fails
State-of-the-art LLMs excel at clinical diagnosis but often fail at Health Economics. When asked to generate treatment plans, they frequently hallucinate costs, ignore local insurance constraints, or suggest financially viable treatments that are practically impossible for the patient.
Medzz-AI… See the full description on the dataset page: https://huggingface.co/datasets/Medzza/oncology-financial-reasoning-india.flask-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Flask
Documentation Data Source Link: https://flask.palletsprojects.com/en/stable/
Data Source License: https://flask.palletsprojects.com/en/stable/license/
Data Source Authors: Pallets Project
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
albumentations-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: albumentations
Documentation Data Source Link: https://github.com/albumentations-team/albumentations?tab=MIT-1-ov-file#readme
Data Source License: https://www.albumentations.ai/docs/
Data Source Authors: Observable AI Benchmarks by Data Agents © 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
weaviate-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: weaviate
Data Source Link: https://weaviate.io/developers/weaviate
Data Source License: https://github.com/weaviate/weaviate/blob/main/LICENSE
Data Source Authors: Weaviate B.V.
AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
Pediatric_medical_reasoningThe generated-pediatric-cases dataset provides a rich collection of synthetic pediatric clinical scenarios designed to support the development and evaluation of reasoning-focused AI models. By offering diverse case studies that include both detailed chain-of-thought reasoning and concise answers, this resource aims to facilitate research in medical decision-making, model interpretability, and educational tools.
This dataset was produced through a data synthesis pipeline powered by the Google… See the full description on the dataset page: https://huggingface.co/datasets/GianlucaMondillo/Pediatric_medical_reasoning.lightning-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: lightning
Documentation Data Source Link: https://lightning.ai/docs/pytorch/stable/
Data Source License: https://github.com/Lightning-AI/pytorch-lightning/blob/master/LICENSE
Data Source Authors: Observable AI Benchmarks by Data Agents © 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
mlflow-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: mlflow
Data Source Link: https://mlflow.org/docs/latest/index.html
Data Source License: https://github.com/mlflow/mlflow/tree/master?tab=Apache-2.0-1-ov-file#readme
Data Source Authors: MLflow Project, a Series of LF Projects, LLC
AI Benchmarks by Data Agents 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
streamlit-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Streamlit
Data Source Link: https://docs.streamlit.io/
Data Source License: https://github.com/streamlit/streamlit/blob/develop/LICENSE
Data Source Authors: Snowflake Inc.
AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
gatsby-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Gatsby.js
Documentation Data Source Link: https://www.gatsbyjs.com/docs/
Data Source License: https://github.com/gatsbyjs/gatsby/blob/master/LICENSE
Data Source Authors: Gatsby
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
go-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Go
Documentation Data Source Link: https://go.dev/doc/
Data Source License: https://go.dev/LICENSE
Data Source Authors: Go Contributors
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
flaml-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: flaml
Documentation Data Source Link: https://microsoft.github.io/FLAML/docs/Getting-Started
Data Source License: https://github.com/microsoft/FLAML?tab=MIT-1-ov-file#readme
Data Source Authors: Observable AI Benchmarks by Data Agents © 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
