datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ethical-Reasoning-in-Mental-Health-v1This repository contains the dataset for the paper EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI.
Overview
Ethical-Reasoning-in-Mental-Health-v1 (EthicsMH) is a carefully curated dataset focused on ethical decision-making scenarios in mental health contexts.This dataset captures the complexity of real-world dilemmas faced by therapists, psychiatrists, and AI systems when navigating critical issues such as confidentiality, autonomy, and bias.
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/UVSKKR/Ethical-Reasoning-in-Mental-Health-v1.ReasoningShield-Dataset
🤗 Dataset Card for ReasoningShield
🛡 1. Dataset Overview
ReasoningShield Dataset is the first comprehensive, well-structured dataset designed to train and evaluate models for detecting hidden safety risks in reasoning traces of Large Reasoning Models (LRMs), spanning 10 risk categories and 3 safety levels. It consists of:
ReasoningShield-Train: 7,000 human-AI annotated (Query… See the full description on the dataset page: https://huggingface.co/datasets/ReasoningShield/ReasoningShield-Dataset.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.agentic-reasoning-benchmark
Agentic & Reasoning Benchmark (ARB) – Expanded
Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning.
Überblick
Eigenschaft
Wert
Anzahl Beispiele
2.550
Kategorien
8
Schwierigkeitsgrade
easy / medium / hard
Formate
CSV + JSON
Reproduzierbarkeit
Generator-Skript (seed=42) enthalten
Lizenz
CC-BY-4.0
Kategorien
Kategorie
Anzahl
Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.DAG-Reasoning-DeepSeek-R1-0528Click here to support our open-source dataset and model releases!
DAG-Reasoning-DeepSeek-R1-0528 is a dataset focused on analysis and reasoning, creating directed acyclic graphs testing the limits of DeepSeek R1 0528's graph-reasoning skills!
This dataset contains:
4.08k synthetically generated prompts to create directed acyclic graphs in response to user input, with all responses generated using DeepSeek R1 0528.
All responses contain a multi-step thinking process to perform effective… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DAG-Reasoning-DeepSeek-R1-0528.multilevel-legal-reasoning
Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations
Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi
Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai.
🧭 Purpose and Scope
The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.finance-reasoning-turkish
Dataset Card for Turkish Advanced Reasoning Dataset (Finance Q&A)
License
This dataset is licensed under the Academic Use Only License. It is intended solely for academic and research purposes. Commercial use is strictly prohibited. For more details, refer to the LICENSE file.
Citation: If you use this dataset in your research, please cite it as follows:
@dataset{turkish_advanced_reasoning_finance_qa,
title = {Turkish Advanced Reasoning Dataset for Finance Q\&A}… See the full description on the dataset page: https://huggingface.co/datasets/emre/finance-reasoning-turkish.finance-reasoning-turkish
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti emre (Davut Emre Tasar, Enes Bulut) tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: emre/finance-reasoning-turkish
🔗 Derleyen Platform: VeriPazarı
Türkçe Gelişmiş Akıl Yürütme Veri Seti (Finans Soru-Cevap)
Lisans
Bu veri seti Sadece Akademik… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/finance-reasoning-turkish.DES-Reasoning-DeepSeek-V3.1Click here to support our open-source dataset and model releases!
DES-Reasoning-DeepSeek-V3.1 is a dataset focused on analysis and reasoning, creating discrete event simulations testing the limits of DeepSeek V3.1's simulation, Python scripting, and analysis skills!
This dataset contains:
4.03k synthetically generated prompts to create discrete event simulations and analysis chat in response to user input, with all responses generated using DeepSeek V3.1.
All responses contain a multi-step… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DES-Reasoning-DeepSeek-V3.1.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains… See the full description on the dataset page: https://huggingface.co/datasets/lucsaint/Deepseek-V4-Reasoning-Code-2500.clinical-authority-reasoning-independence-v0.1Clinical Decision–Constraint Integrity v0.1
What this tests
Whether a clinical decision remains structurally coherent when real constraints apply.
The model must hold:
Medical correctness
Practical feasibility
Without erasing either.
Failure modes
constraint_erasedThe decision ignores or deletes the constraint
false_resolutionThe response pretends the conflict does not exist
coherent_tradeoffThe response names limits and adapts without distortion
How it works
Decision context defines the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-authority-reasoning-independence-v0.1.oncology-financial-reasoning-india
🩺 Medzz-AI: Oncology & Financial Reasoning (India)
Status: Active | Context: Indian Healthcare | Focus: Clinical + Economic Logic
👋 The Problem: Why Current Medical AI Fails
State-of-the-art LLMs excel at clinical diagnosis but often fail at Health Economics. When asked to generate treatment plans, they frequently hallucinate costs, ignore local insurance constraints, or suggest financially viable treatments that are practically impossible for the patient.
Medzz-AI… See the full description on the dataset page: https://huggingface.co/datasets/Medzza/oncology-financial-reasoning-india.Drone-flight-monitoring-reasoning-SFT
Drone-flight-monitoring-reasoning-SFT
Dataset Description
本数据集是一个专注于无人机飞行安全领域的中文问答数据集,采用了Chain-of-Thought (CoT) 的格式。它旨在用于练习大语言模型的微调训练,使其能够模拟专家思考过程,并针对无人机安全相关问题生成包含推理步骤的结构化回答。微调后模型见(GabrielCheng/Deepseek-r1-finetuned-drone-safty) 。
本数据集是基于 Hugging Face 平台上的 skylink-drone-cot-datasets (pohsjxx/default-domain-cot-dataset) 进行处理和衍生的。
Dataset Structure / Data Fields
数据集中的每个样本包含以下字段:
Question (string): 关于无人机飞行安全或风险相关的问题。
Reasoning (string): 模拟模型的推理过程。
Answer… See the full description on the dataset page: https://huggingface.co/datasets/GabrielCheng/Drone-flight-monitoring-reasoning-SFT.pashto-opus-5k-reasoning-max
🚀 Pashto OPUS 5K Reasoning Max
This dataset is a high-quality collection of 5,000 reasoning-focused pairs, derived from the OPUS corpus and enhanced for Pashto Language Models. It is specifically curated to push the boundaries of "Chain-of-Thought" (CoT) and logical deduction in the Pashto language.
🌟 Overview
While standard OPUS data is often used for simple translation, Pashto-OPUS-5K-Reasoning-Max takes it a step further by focusing on complex instructions and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-opus-5k-reasoning-max.nvidia-nemotron-model-reasoning-dataset-turkish
Nemotron Reasoning Challenge - Turkish
Turkish translation of the training data from NVIDIA's Nemotron Model Reasoning Challenge
Each row is a reasoning puzzle framed in an "Alice's Wonderland" setting. Given a few input/output examples, the model needs to figure out the hidden rule and apply it to a new input.
Category
Rows
Description
bit
1602
Hidden bit manipulation rule on 8-bit binary numbers
grav
1597
Falling distance with a modified gravitational constant… See the full description on the dataset page: https://huggingface.co/datasets/mramazan/nvidia-nemotron-model-reasoning-dataset-turkish.Reasoning-Hypothesis-Corpus
Reasoning-Hypothesis-Corpus Dataset
Overview
The Reasoning-Hypothesis-Corpus is a private dataset designed for tasks involving reasoning and hypothesis evaluation. The dataset consists of pairs of premises and corresponding hypotheses. Each entry aims to help models understand and reason about the relationship between textual descriptions.
Modality: Text
Format: CSV
Size: <1K rows
License: Apache 2.0
Dataset Structure
Split
Train:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Reasoning-Hypothesis-Corpus.Italian_reasoning_dataset
Quiz ed enigmi logici in italiano - Synthetic Dataset (Partial Data)
Note: Dataset generation aimed for 1100 rows, but 985 rows were successfully produced. This may be due to model output characteristics, filtering of incomplete items, or automatic correction of JSON key names.
This dataset was generated using the Synthetic Dataset Generator powered by Gemini AI.
Topic: Quiz ed enigmi logici in italiano
Field 1: domanda e risposta corretta delimitati da tag
Field 2: Il pensiero… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/Italian_reasoning_dataset.reasoning-persona-dataset_test
Reasoning + Persona SFT Dataset
Columns: instruction, input, output, persona, reasoning_summaryUse: Supervised fine-tuning for cinematic/storytelling or creative-director style outputs.
Schema
instruction (str)
input (str)
output (str)
persona (str)
reasoning_summary (str, brief rationale cue)
Citation
Author: saravan
reasoning-prompt-ko
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/Mediaapel/reasoning-prompt-ko.schema_cot_reasoning
⊙ Prompt Programs for Agentic Reasoning
Programmable task-dependent COTs for agentic reasoning.
A 100-row seed dataset for programmable cognition.
Each row defines:
Prompt template + input binding + explanation + task-dependent reasoning program
Pipeline:
intake → binding → procedure → output
Schema
Column
Meaning
ID
Stable row ID
Name
Task name
Prompt
Prompt template using {{VARIABLE}}
Expression
Input binding using $.path
Explanation
Binding… See the full description on the dataset page: https://huggingface.co/datasets/bitwikiorg/schema_cot_reasoning.
