CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MiniMaxAI /SynLogic SynLogic Dataset SynLogic is a comprehensive synthetic logical reasoning dataset designed to enhance logical reasoning capabilities in Large Language Models (LLMs) through reinforcement learning with verifiable rewards. 🐙 GitHub Repo: https://github.com/MiniMax-AI/SynLogic 📜 Paper (arXiv): https://arxiv.org/abs/2505.19641 Dataset Description SynLogic contains 35 diverse logical reasoning tasks with automatic verification capabilities, making it ideal for… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/SynLogic.texttext-generation10K<n<100K106 likes1.8k downloads1y agoHugging Face02PursuitOfDataScience /MiniMax-M2.1-Mixture-of-Thoughts MiniMax-M2.1 Mixture of Thoughts This dataset contains responses generated by MiniMax-M2.1 for user questions from the open-r1/Mixture-of-Thoughts dataset. Dataset Description The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation. Metric Value Examples 349,317 Total Tokens 4,052,592,552 Avg Tokens/Example 11,601 Source Dataset Name:… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts.tabulartext-generation100K<n<1M2 likes1.2k downloads9mo agoHugging Face03oakmindai /minimax_h3_avatar_500 Watch the full 500-video showcase on YouTube MiniMax H3 Avatar 500 An image-to-video dataset pairing reference avatar images with detailed generation prompts and generated avatar videos. This release contains 500 curated examples in both a browsable raw layout and a typed Hugging Face dataset. Version 1.0 · Released August 14, 2026 Dataset contents Each example contains: A 1024 × 1024 reference avatar image A detailed English generation prompt A generated 640 ×… See the full description on the dataset page: https://huggingface.co/datasets/oakmindai/minimax_h3_avatar_500.imagen<1K3 likes826 downloads1mo agoHugging Face04AlienKevin /SWE-smith-rs-minimax-m2.5-trajectories Trajectories Dataset Top-level fields: messages instance_id resolved model traj_id patch Generated at: 2026-02-27 00:04:28Z Rows: 5251 Shards: 21 Skipped runs (missing/corrupt trajectory): 60 text1K<n<10K3 likes547 downloads7mo agoHugging Face05MiniMaxAI /role-play-bench Role-play Benchmark A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios. Dataset Summary Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?".… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/role-play-bench.tabulartext-generation1K<n<10K151 likes454 downloads8mo agoHugging Face06Nilaksh404 /minimax-m2documentn<1K0 likes370 downloads10mo agoHugging Face07MiniMaxAI /VIBE VIBE: Visual & Interactive Benchmark for Execution in Application Development [English] | 中文 🌟 Overview VIBE (Visual & Interactive Benchmark for Execution) sets a new standard for evaluating Large Language Models (LLMs) in full-stack software engineering. Moving beyond recent benchmarks that rely on static screenshots or rigid workflow snapshots to assess application development, VIBE pioneers the Agent-as-a-Verifier (AaaV) paradigm to assess the true "0-to-1" capability… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/VIBE.texttext-generationn<1K279 likes272 downloads9mo agoHugging Face08open-athena /nemotron-gym-instruction-following-structured-minimax-m27-131k-tracestext1K<n<10K0 likes196 downloads4mo agoHugging Face09open-athena /MiniMax-M2.7-stackexchange-tezos-sandboxes-maxeps-32k-juptext1K<n<10K0 likes125 downloads4mo agoHugging Face10AmanPriyanshu /reasoning-sft-minimax-microsoft-orca-agentinstruct-1M-v1 MiniMax-M2.5 Reasoning SFT (Orca AgentInstruct 1M v1) Reasoning SFT dataset generated by MiniMaxAI/MiniMax-M2.5 on prompts from the Stratified K-Means Diverse Instruction-Following 100K-1M dataset (Orca AgentInstruct subset). Format Each row has three columns: input — list of dicts [{"role": "...", "content": "..."}, ...] (conversation turns) response — model-generated response with <think> reasoning block source — task category (creative_content, text_modification, rc… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-minimax-microsoft-orca-agentinstruct-1M-v1.texttext-generation100K<n<1M1 likes117 downloads6mo agoHugging Face11VINAY-UMRETHE /Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-Highgated Distill This is a multi-source curated instruction and reasoning dataset specifically for training and distilling large language models (LLMs) to exhibit advanced Chain-of-Thought (CoT), Agentic, Mathematical and Coding capabilities. It aggregates high-quality outputs from frontier models into messages ChatML format. Dataset Structure The dataset contains a total of 70.2K examples, split into three subsets based on the presence of visible reasoning… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High.texttext-generation100K<n<1M14 likes95 downloads3mo agoHugging Face12OpenMed /Medical-Reasoning-SFT-MiniMax-M2.1 Medical-Reasoning-SFT-MiniMax-M2.1 A large-scale medical reasoning dataset generated using MiniMaxAI/MiniMax-M2.1, containing over 204,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model MiniMaxAI/MiniMax-M2.1 Total Samples 204,773 Samples with Reasoning 204,773 (100%) Estimated Tokens ~621 Million Content Tokens ~344 Million Reasoning Tokens ~277 Million Language English… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-MiniMax-M2.1.texttext-generation100K<n<1M9 likes87 downloads8mo agoHugging Face13faezeb /minimax21-completetext100K<n<1M0 likes87 downloads8mo agoHugging Face14AmanPriyanshu /reasoning-sft-minimax-stratified-kmeans-diverse-reasoning-842K-only MiniMax-M2.5 Reasoning SFT (Stratified K-Means Diverse Reasoning 1M) Reasoning SFT dataset generated by MiniMaxAI/MiniMax-M2.5 on prompts from the Stratified K-Means Diverse Reasoning 100K-1M dataset. Format Each row has three columns: input — list of dicts [{"role": "...", "content": "..."}, ...] (conversation turns) response — model-generated response with <think> reasoning block source — task category (math, code, science, chat, safety) Generation Model:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-minimax-stratified-kmeans-diverse-reasoning-842K-only.texttext-generation100K<n<1M0 likes84 downloads6mo agoHugging Face15faezeb /minimax21-deduped-featurestabular100K<n<1M0 likes81 downloads7mo agoHugging Face16minimaxir /mtg-embeddings Dataset Card for Dataset Name Text embeddings of all Magic: The Gathering card until Aetherdrift (2024-02-14). The text embeddings are centered around card mechanics (i.e. no flavor text/card art embeddings) in order to identify similar cards mathematically. This dataset also includes zero-mean-centered 2D UMAP coordinates for all the cards, in columns x_2d and y_2d. Dataset Details How The Embeddings Were Created Using the data exports from MTGJSON, the data… See the full description on the dataset page: https://huggingface.co/datasets/minimaxir/mtg-embeddings.tabular10K<n<100K6 likes77 downloads2y agoHugging Face17minimaxir /llm-blueberrytabular1K<n<10K0 likes67 downloads1y agoHugging Face18faezeb /dolci-base-minimax-m2-completions-featurestabular100K<n<1M0 likes67 downloads7mo agoHugging Face19open-athena /selfinstruct-naive-sandboxes-2-verified-minimax-m27-131k-tracestext1K<n<10K0 likes63 downloads3mo agoHugging Face20open-athena /llm-verifier-freelancer-minimax-m27-131k-tracestext1K<n<10K0 likes62 downloads4mo agoHugging Face21penfever /swe_rebench_patched_oracle-minimax-m27-131k-traces_chunk0textn<1K0 likes56 downloads3mo agoHugging Face22DCAgent2 /DCAgent2_terminal_bench_2_laion_MiniMax-M2-freelancer-32ep-32k-reasoning_2025113e0550cftextn<1K0 likes54 downloads10mo agoHugging Face23Madras1 /minimax-m2.5-code-distilled-14k MiniMax M2.5 Code Distillation Dataset A synthetic code generation dataset created by distilling **MiniMax-M2.5. Each example contains a Python coding problem, the model's chain-of-thought reasoning, and a verified correct solution that passes automated test execution. Key Features Execution-verified: Every solution was executed against test cases in a sandboxed subprocess. Only solutions that passed all tests are included. Chain-of-thought reasoning: Each example… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/minimax-m2.5-code-distilled-14k.texttext-generation10K<n<100K15 likes54 downloads4mo agoHugging Face24DCAgent2 /DCAgent2_terminal_bench_2_laion_MiniMax-M2-freelancer-32ep-32k_20251129_083530textn<1K0 likes52 downloads10mo agoHugging Face25open-athena /nemotron-gym-agent-workplace-v2-minimax-m27-131k-tracestextn<1K0 likes49 downloads4mo agoHugging Face26open-athena /nemotron-gym-identity-following-v2-minimax-m27-131k-tracestext10K<n<100K0 likes49 downloads3mo agoHugging Face27open-athena /swegym-tasks-patched-validated-v5-minimax-m27-131k-tracestext1K<n<10K0 likes49 downloads3mo agoHugging Face28open-athena /exp_rpt_nemotron-junit-minimax-m27-131k-tracestext1K<n<10K0 likes48 downloads4mo agoHugging Face29open-athena /swesmith-oracle-filtered-minimax-m27-131k-tracestext10K<n<100K0 likes48 downloads3mo agoHugging Face30DCAgent2 /DCAgent_dev_set_71_tasks_laion_MiniMax-M2-freelancer-32ep-32k-reasoning_20251128_092618textn<1K0 likes47 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.