datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MolmoAct-Midtraining-Mixture
MolmoAct - Midtraining Mixture
Data Mixture used for MolmoAct Midtraining. Contains MolmoAct Dataset formulated as Action Reasoning Data.
MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique manipulation tasks in both home and tabletop environments. It has… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Midtraining-Mixture.tulu-3-sft-mixture
Tulu 3 SFT Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
The Tulu 3 SFT mixture was used to train the Tulu 3 series of models.
It contains 939,344 samples from the following sets:
CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024)
FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-mixture.Mixture-of-Thoughts
Dataset summary
Mixture-of-Thoughts is a curated dataset of 350k verified reasoning traces distilled from DeepSeek-R1. The dataset spans tasks in mathematics, coding, and science, and is designed to teach language models to reason step-by-step. It was used in the Open R1 project to train OpenR1-Distill-7B, an SFT model that replicates the reasoning capabilities of deepseek-ai/DeepSeek-R1-Distill-Qwen-7B from the same base model.
To load the dataset, run:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/Mixture-of-Thoughts.MolmoAct-Pretraining-Mixture
MolmoAct - Pretraining Mixture
Data Mixture used for MolmoAct Pretraining. Contains a subset of OXE formulated as Action Reasoning Data along with auxiliary robot data and link to Multimodal Web data.
MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Pretraining-Mixture.llama-3.1-tulu-3-8b-preference-mixture
Tulu 3 8B Preference Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This mix is made up from the following preference datasets:
https://huggingface.co/datasets/allenai/tulu-3-sft-reused-off-policy
https://huggingface.co/datasets/allenai/tulu-3-sft-reused-on-policy-8b… See the full description on the dataset page: https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-8b-preference-mixture.tulu-v2-sft-mixture
Dataset Card for Tulu V2 Mix
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
Tulu is a series of language models that are trained to act as helpful assistants.
The dataset consists of a mix of :
FLAN (Apache 2.0): We use 50,000 examples sampled from FLAN v2. To emphasize CoT-style reasoning, we sample another 50,000 examples… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture.tulu-3-sft-olmo-2-mixture-0225Used to train OLMo 2 32B. From the blog post:
Filtered out instructions from the SFT dataset and the chosen responses of the preference data that included mentions of a date cutoff from the synthetic data generation process. This resulted in a new version of the instruction dataset, Tulu 3 SFT Mixture 0225, and preference dataset, OLMo-2-32B-pref-mix-0325.
We use majority voting to improve the quality of answers to our synthetic math questions. For our Persona MATH and Grade School Math… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture-0225.MiniMax-M2.1-Mixture-of-Thoughts
MiniMax-M2.1 Mixture of Thoughts
This dataset contains responses generated by MiniMax-M2.1 for user questions from the open-r1/Mixture-of-Thoughts dataset.
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Examples
349,317
Total Tokens
4,052,592,552
Avg Tokens/Example
11,601
Source Dataset
Name:… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts.stage3-final-mixtureindic-oss-mixture-cpt-10btulu-3-sft-olmo-2-mixtureNote that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
The OLMo v2 SFT mixture was used to train the OLMo models.
It contains 939,344 samples from the following sets:
CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024)
FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et al., 2023)
No Robots (CC-BY-NC-4.0), 9,500… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture.stage3-final-mixture-cot50
Stage 3 Final Mixture — 50% CoT Compression
This is a deterministic capability-preserving rewrite of
leonli66/stage3-final-mixture for LCLM Stage-3 post-training.
Only the reasoning_data and dolci_think subsets change. Their
compression_prompt is the ordinary prompt. A deterministic 50% arm keeps
the complete assistant target as ordinary SFT; the other arm wraps the inferred
reasoning prefix in <|memory_start|>...<|memory_end|> while keeping the final
answer trainable. All… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-final-mixture-cot50.ODA-Mixture-500k
ODA-Mixture-500k
ODA-Mixture-500k is a large-scale general-purpose post-training dataset curated from top-performing open corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination.
🧠 Dataset Summary
Domain: General-purpose(e.g., Math, Code, Reasoning, General).
Format: Problem → Solution (reasoning trace) → Final answer.
Scale (selected training set): ~500K samples.
Goal: Achieve maximum general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Mixture-500k.ohlc_1d_mixture
Macroeconomic & S&P 500 Yahoo Finance Dataset
This repository contains a comprehensive historical dataset for 936 financial instruments, including S&P 500 components, broad market indices, commodities, currencies, and macroeconomic indicators. The data is programmatically extracted from the Yahoo Finance API, cleaned, and normalized for use in quantitative modeling and machine learning.
Dataset Hub: bguzzo2k/ohlc_1d_mixture
Repository Structure
1d/: Raw daily OHLCV… See the full description on the dataset page: https://huggingface.co/datasets/bguzzo2k/ohlc_1d_mixture.mixturerlhflow_mixture_with_math_del_systemlf3-mixture-data
lf3-mixture-data
Citation and attribution
This dataset repository is maintained by Gökay Aydoğan. If you reference this repository in academic work, please cite it as follows and also cite the upstream models, datasets, or projects it builds upon.
@dataset{aydogan2026lf3_mixture_data,
author = {Aydoğan, Gökay},
title = {{lf3-mixture-data}},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/gokaygokay/lf3-mixture-data}… See the full description on the dataset page: https://huggingface.co/datasets/gokaygokay/lf3-mixture-data.stage3-mixture-cleaneurospeech-bg-diar-mixtures
⚠️ DEPRECATED — use v2
This dataset contains a shortcut that lets a model infer the number of
speakers without listening to the audio.
Each speaker was given a fixed 4 turns, so session duration is a direct
function of speaker count. Measured on this data:
1-spk median 18.2 s range 11.1-25.2
2-spk 30.4 s 23.0-36.6
3-spk 41.6 s 21.0-54.5
4-spk 54.6 s 38.0-73.0
The 2-speaker and 4-speaker ranges do not overlap — 2-spk tops out… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/eurospeech-bg-diar-mixtures.Franka2_pour_stir_and_shake_mixture_0330zip2zip-plus-mixture-partitioned
Zip2Zip Plus Mixture Partitioned
This dataset is a partitioned pretraining-data mixture built for zip2zip language-model pretraining.
The mixture is byte-balanced across four top-level domains:
Domain
Source
Target byte ratio
General
HuggingFaceFW/fineweb-edu, sample-100BT
50%
Code
bigcode/the-stack-dedup
20%
Math
HuggingFaceTB/finemath, finemath-3plus
10%
Multilingual
epfml/FineWeb2-HQ, 20 language subsets
20%
The uploaded layout is partitioned by source… See the full description on the dataset page: https://huggingface.co/datasets/mxxsc/zip2zip-plus-mixture-partitioned.SmolLM-lmsys-mixturescolsmol-mixturedetails_jsfs11__MixtureofMerges-MoE-4x7b-v4
Dataset Card for Evaluation run of jsfs11/MixtureofMerges-MoE-4x7b-v4
Dataset automatically created during the evaluation run of model jsfs11/MixtureofMerges-MoE-4x7b-v4.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_jsfs11__MixtureofMerges-MoE-4x7b-v4.qsharp-full-mixture-1.5b-filtered-with-labelsmixture_ami_synthetic_bigODA-Mixture-500k-copy-58
ODA-Mixture-500k
ODA-Mixture-500k is a large-scale general-purpose post-training dataset curated from top-performing open corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination.
🧠 Dataset Summary
Domain: General-purpose(e.g., Math, Code, Reasoning, General).
Format: Problem → Solution (reasoning trace) → Final answer.
Scale (selected training set): ~500K samples.
Goal: Achieve maximum general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/Thesho1/ODA-Mixture-500k-copy-58.preference_dataset_mixture2_and_safe_pku
Copy from https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question… See the full description on the dataset page: https://huggingface.co/datasets/OpenRLHF/preference_dataset_mixture2_and_safe_pku.ODA-Mixture-100k
ODA-Mixture-100k
ODA-Mixture-100k is a compact general-purpose post-training dataset curated from top-performing open corpora (selected via the *OpenDataArena* leaderboard) and refined through deduplication, benchmark decontamination.
🧠 Dataset Summary
Domain: General-purpose(e.g., Math, Code, Reasoning, General).
Format: Problem → Solution (reasoning trace) → Final answer.
Scale (selected training set): ~100K samples.
Goal: Achieve significant general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Mixture-100k.ODA-Mixture-500k-copy-53
ODA-Mixture-500k
ODA-Mixture-500k is a large-scale general-purpose post-training dataset curated from top-performing open corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination.
🧠 Dataset Summary
Domain: General-purpose(e.g., Math, Code, Reasoning, General).
Format: Problem → Solution (reasoning trace) → Final answer.
Scale (selected training set): ~500K samples.
Goal: Achieve maximum general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/Thesho1/ODA-Mixture-500k-copy-53.
