datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MiniMax-M2.1-Mixture-of-Thoughts
MiniMax-M2.1 Mixture of Thoughts
This dataset contains responses generated by MiniMax-M2.1 for user questions from the open-r1/Mixture-of-Thoughts dataset.
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Examples
349,317
Total Tokens
4,052,592,552
Avg Tokens/Example
11,601
Source Dataset
Name:… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts.tulu_sft_mixture_preprocessed
Tulu SFT Mixture Preprocessed
This dataset was created by preprocessing the
allenai/tulu-3-sft-mixture
dataset for single-turn supervised fine-tuning.
The preprocessing keeps English user -> assistant examples from the selected
Tulu sources, applies length filtering with the official
Qwen/Qwen3.5-4B-Base chat template, and removes exact and near duplicates.
The resulting train split contains 151,292 examples with a maximum sequence
length of 7,168 tokens.
Each row contains the… See the full description on the dataset page: https://huggingface.co/datasets/HwanChang0106/tulu_sft_mixture_preprocessed.mixture-of-experts-papers
Mixture of Experts Papers — FineSet
A research-paper dataset on Mixture of Experts Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Mixture of Experts Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/mixture-of-experts-papers.mixture-then-select-selections
Frozen Qwen3-8B Selections
This dataset contains the exact selection metadata and pool indices used for
a frozen cross-scale data-selection experiment. The subsets were selected
using Qwen3-8B-derived information and can be transferred unchanged to a
tokenizer-compatible larger target model.
The companion training and evaluation code is:
https://github.com/submissionpaper1234/mixture-then-select-reproducibility
Critical Interpretation
The selected instruction text… See the full description on the dataset page: https://huggingface.co/datasets/submissionpaper1234/mixture-then-select-selections.
