datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SynLogic
SynLogic Dataset
SynLogic is a comprehensive synthetic logical reasoning dataset designed to enhance logical reasoning capabilities in Large Language Models (LLMs) through reinforcement learning with verifiable rewards.
🐙 GitHub Repo: https://github.com/MiniMax-AI/SynLogic
📜 Paper (arXiv): https://arxiv.org/abs/2505.19641
Dataset Description
SynLogic contains 35 diverse logical reasoning tasks with automatic verification capabilities, making it ideal for… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/SynLogic.MiniMax-M2.1-Mixture-of-Thoughts
MiniMax-M2.1 Mixture of Thoughts
This dataset contains responses generated by MiniMax-M2.1 for user questions from the open-r1/Mixture-of-Thoughts dataset.
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Examples
349,317
Total Tokens
4,052,592,552
Avg Tokens/Example
11,601
Source Dataset
Name:… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts.role-play-bench
Role-play Benchmark
A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios.
Dataset Summary
Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?".… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/role-play-bench.VIBE
VIBE: Visual & Interactive Benchmark for Execution in Application Development
[English] | 中文
🌟 Overview
VIBE (Visual & Interactive Benchmark for Execution) sets a new standard for evaluating Large Language Models (LLMs) in full-stack software engineering. Moving beyond recent benchmarks that rely on static screenshots or rigid workflow snapshots to assess application development, VIBE pioneers the Agent-as-a-Verifier (AaaV) paradigm to assess the true "0-to-1" capability… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/VIBE.Medical-Reasoning-SFT-MiniMax-M2.1
Medical-Reasoning-SFT-MiniMax-M2.1
A large-scale medical reasoning dataset generated using MiniMaxAI/MiniMax-M2.1, containing over 204,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
MiniMaxAI/MiniMax-M2.1
Total Samples
204,773
Samples with Reasoning
204,773 (100%)
Estimated Tokens
~621 Million
Content Tokens
~344 Million
Reasoning Tokens
~277 Million
Language
English… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-MiniMax-M2.1.Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High
Distill
This is a multi-source curated instruction and reasoning dataset specifically for training and distilling large language models (LLMs) to exhibit advanced Chain-of-Thought (CoT), Agentic, Mathematical and Coding capabilities. It aggregates high-quality outputs from frontier models into messages ChatML format.
Dataset Structure
The dataset contains a total of 70.2K examples, split into three subsets based on the presence of visible reasoning… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High.reasoning-sft-minimax-microsoft-orca-agentinstruct-1M-v1
MiniMax-M2.5 Reasoning SFT (Orca AgentInstruct 1M v1)
Reasoning SFT dataset generated by MiniMaxAI/MiniMax-M2.5 on prompts from the Stratified K-Means Diverse Instruction-Following 100K-1M dataset (Orca AgentInstruct subset).
Format
Each row has three columns:
input — list of dicts [{"role": "...", "content": "..."}, ...] (conversation turns)
response — model-generated response with <think> reasoning block
source — task category (creative_content, text_modification, rc… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-minimax-microsoft-orca-agentinstruct-1M-v1.reasoning-sft-minimax-stratified-kmeans-diverse-reasoning-842K-only
MiniMax-M2.5 Reasoning SFT (Stratified K-Means Diverse Reasoning 1M)
Reasoning SFT dataset generated by MiniMaxAI/MiniMax-M2.5 on prompts from the Stratified K-Means Diverse Reasoning 100K-1M dataset.
Format
Each row has three columns:
input — list of dicts [{"role": "...", "content": "..."}, ...] (conversation turns)
response — model-generated response with <think> reasoning block
source — task category (math, code, science, chat, safety)
Generation
Model:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-minimax-stratified-kmeans-diverse-reasoning-842K-only.minimax-m2.5-code-distilled-14k
MiniMax M2.5 Code Distillation Dataset
A synthetic code generation dataset created by distilling **MiniMax-M2.5. Each example contains a Python coding problem, the model's chain-of-thought reasoning, and a verified correct solution that passes automated test execution.
Key Features
Execution-verified: Every solution was executed against test cases in a sandboxed subprocess. Only solutions that passed all tests are included.
Chain-of-thought reasoning: Each example… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/minimax-m2.5-code-distilled-14k.MiniMaxAI-VIBE
VIBE: Visual & Interactive Benchmark for Execution in Application Development
[English] | 中文
🌟 Overview
VIBE (Visual & Interactive Benchmark for Execution) sets a new standard for evaluating Large Language Models (LLMs) in full-stack software engineering. Moving beyond recent benchmarks that rely on static screenshots or rigid workflow snapshots to assess application development, VIBE pioneers the Agent-as-a-Verifier (AaaV) paradigm to assess the true "0-to-1" capability… See the full description on the dataset page: https://huggingface.co/datasets/John1604/MiniMaxAI-VIBE.
