datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MiniMax-M2.1-Mixture-of-Thoughts
MiniMax-M2.1 Mixture of Thoughts
This dataset contains responses generated by MiniMax-M2.1 for user questions from the open-r1/Mixture-of-Thoughts dataset.
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Examples
349,317
Total Tokens
4,052,592,552
Avg Tokens/Example
11,601
Source Dataset
Name:… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts.ohlc_1d_mixture
Macroeconomic & S&P 500 Yahoo Finance Dataset
This repository contains a comprehensive historical dataset for 936 financial instruments, including S&P 500 components, broad market indices, commodities, currencies, and macroeconomic indicators. The data is programmatically extracted from the Yahoo Finance API, cleaned, and normalized for use in quantitative modeling and machine learning.
Dataset Hub: bguzzo2k/ohlc_1d_mixture
Repository Structure
1d/: Raw daily OHLCV… See the full description on the dataset page: https://huggingface.co/datasets/bguzzo2k/ohlc_1d_mixture.eurospeech-bg-diar-mixtures
⚠️ DEPRECATED — use v2
This dataset contains a shortcut that lets a model infer the number of
speakers without listening to the audio.
Each speaker was given a fixed 4 turns, so session duration is a direct
function of speaker count. Measured on this data:
1-spk median 18.2 s range 11.1-25.2
2-spk 30.4 s 23.0-36.6
3-spk 41.6 s 21.0-54.5
4-spk 54.6 s 38.0-73.0
The 2-speaker and 4-speaker ranges do not overlap — 2-spk tops out… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/eurospeech-bg-diar-mixtures.details_jsfs11__MixtureofMerges-MoE-4x7b-v4
Dataset Card for Evaluation run of jsfs11/MixtureofMerges-MoE-4x7b-v4
Dataset automatically created during the evaluation run of model jsfs11/MixtureofMerges-MoE-4x7b-v4.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_jsfs11__MixtureofMerges-MoE-4x7b-v4.preference_dataset_mixture2_and_safe_pku
Copy from https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question… See the full description on the dataset page: https://huggingface.co/datasets/OpenRLHF/preference_dataset_mixture2_and_safe_pku.open-hermes-2.5-sft-mixture-llama3-inference-retrieval-tokenspreference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question with this reward model and also any question about reward modeling, feel free to drop me an… See the full description on the dataset page: https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku.llama-3.1-tulu-3-405b-preference-mixture
Llama 3.1 Tulu 3 405B Preference Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This preference mixture used for DPO on our the Llama 3.1 Tulu 3 405B SFT checkpoint to obtain Llama 3.1 Tulu 3 405B DPO.
It contains 360,924 generation pairs obtained using the following models:
Mistral 7B Instruct v0.2 (Apache 2.0)… See the full description on the dataset page: https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-405b-preference-mixture.stage3-final-mixture-20260910-packed
Stage 3 final packed training mixture
Version
Packed sequences
Original examples retained
Parquet files
16,384
2,278,921
21,463,686
4,452
32,768
1,277,010
21,669,429
2,495
Each sequence contains only one category: agent, reasoning or other. Completed
sequences are globally shuffled together using PCG64 seed 20260910; both native
and expanded agents are included. All shuffle positions and serialized payload
preservation were checked. Original examples overlap… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-final-mixture-20260910-packed.preference_dataset_mixture2_and_safe_pku150k
Dataset Card for "preference_dataset_mixture2_and_safe_pku150k"
More Information needed
preference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question with this reward model and also any question about reward modeling, feel free to drop me an… See the full description on the dataset page: https://huggingface.co/datasets/living-box/preference_dataset_mixture2_and_safe_pku.tulu-v2-sft-mixture-first-stage-classifier-outputsmixture-binarized
Copy from https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question… See the full description on the dataset page: https://huggingface.co/datasets/gohsyi/mixture-binarized.tulu-v2-sft-mixture-first-stage-classifier-outputs-v1tulu_sft_mixture_preprocessed
Tulu SFT Mixture Preprocessed
This dataset was created by preprocessing the
allenai/tulu-3-sft-mixture
dataset for single-turn supervised fine-tuning.
The preprocessing keeps English user -> assistant examples from the selected
Tulu sources, applies length filtering with the official
Qwen/Qwen3.5-4B-Base chat template, and removes exact and near duplicates.
The resulting train split contains 151,292 examples with a maximum sequence
length of 7,168 tokens.
Each row contains the… See the full description on the dataset page: https://huggingface.co/datasets/HwanChang0106/tulu_sft_mixture_preprocessed.code_rlvr_mixture_dpomultidomain-planner-mixture-v6-curriculum-2000-glm5-trim-4-reversedpreference_dataset_mixture
Dataset Card for "preference_dataset_mixture"
The dataset used to train weqweasdas/RM-Gemma-7B . See the model page for details.
multidomain-planner-mixture-v7nanochat-four-capability-mixture-v3-5x-20260714-r1
Nanochat Capability Mixture
This corpus mixes depo, brevo, world_state, tool_routing with deterministic documents balancing. Training and validation rows only are included; source test rows remain excluded. Exact source manifests are retained under provenance/, and source-aware validation can replay every selected row against the authoritative datasets.
llama-3.1-tulu-3-405b-preference-mixture-filter-datecutofftulu-v2-sft-mixture-second-stage-classifier-v1-stage2-example-based-factual-info-outputsmultidomain-planner-mixture-v6tulu-v2-sft-mixture-llama3-inference-syn-knowledge-outputs-v1llama-3.1-tulu-3-8b-preference-mixture-safety-reduced
llama-3.1-tulu-3-8b-preference-mixture-safety-reduced
Safety labeling provenance
Every example in the input dataset iamjanvijay/llama-3.1-tulu-3-8b-preference-mixture was safety-labeled using:
Default labeling model: Qwen/Qwen3-32B
Default decoding params: temperature=0.85, top_p=0.8, top_k=20, min_p=0.0
Judge prompt
You are a binary classifier for detecting broad safety/alignment behavior in assistant responses.
Task:
Given a conversation, determine whether… See the full description on the dataset page: https://huggingface.co/datasets/iamjanvijay/llama-3.1-tulu-3-8b-preference-mixture-safety-reduced.multidomain-planner-mixture-v6-curriculum-2000-glm5-trimtulu-3-sft-mixture-split-seed-1337-train-qwen-tokenized-lenstulu-3-sft-mixture-split-seed-1337multidomain-planner-mixture-v5multidomain-planner-mixture-v6-curriculum-2000-glm5
