CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /MolmoAct-Midtraining-Mixture MolmoAct - Midtraining Mixture Data Mixture used for MolmoAct Midtraining. Contains MolmoAct Dataset formulated as Action Reasoning Data. MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique manipulation tasks in both home and tabletop environments. It has… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Midtraining-Mixture.imagerobotics1M<n<10M6 likes66k downloads1y agoHugging Face02allenai /tulu-3-sft-mixture Tulu 3 SFT Mixture Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. The Tulu 3 SFT mixture was used to train the Tulu 3 series of models. It contains 939,344 samples from the following sets: CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024) FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-mixture.textother100K<n<1M265 likes61k downloads2y agoHugging Face03open-r1 /Mixture-of-Thoughts Dataset summary Mixture-of-Thoughts is a curated dataset of 350k verified reasoning traces distilled from DeepSeek-R1. The dataset spans tasks in mathematics, coding, and science, and is designed to teach language models to reason step-by-step. It was used in the Open R1 project to train OpenR1-Distill-7B, an SFT model that replicates the reasoning capabilities of deepseek-ai/DeepSeek-R1-Distill-Qwen-7B from the same base model. To load the dataset, run: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/Mixture-of-Thoughts.texttext-generation100K<n<1M333 likes15k downloads1y agoHugging Face04allenai /MolmoAct-Pretraining-Mixture MolmoAct - Pretraining Mixture Data Mixture used for MolmoAct Pretraining. Contains a subset of OXE formulated as Action Reasoning Data along with auxiliary robot data and link to Multimodal Web data. MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Pretraining-Mixture.imagerobotics10M<n<100M14 likes11k downloads1y agoHugging Face05allenai /llama-3.1-tulu-3-8b-preference-mixture Tulu 3 8B Preference Mixture Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. This mix is made up from the following preference datasets: https://huggingface.co/datasets/allenai/tulu-3-sft-reused-off-policy https://huggingface.co/datasets/allenai/tulu-3-sft-reused-on-policy-8b… See the full description on the dataset page: https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-8b-preference-mixture.text100K<n<1M27 likes2.9k downloads2y agoHugging Face06allenai /tulu-v2-sft-mixture Dataset Card for Tulu V2 Mix Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact. Tulu is a series of language models that are trained to act as helpful assistants. The dataset consists of a mix of : FLAN (Apache 2.0): We use 50,000 examples sampled from FLAN v2. To emphasize CoT-style reasoning, we sample another 50,000 examples… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture.textquestion-answering100K<n<1M138 likes2.7k downloads2y agoHugging Face07allenai /tulu-3-sft-olmo-2-mixture-0225Used to train OLMo 2 32B. From the blog post: Filtered out instructions from the SFT dataset and the chosen responses of the preference data that included mentions of a date cutoff from the synthetic data generation process. This resulted in a new version of the instruction dataset, Tulu 3 SFT Mixture 0225, and preference dataset, OLMo-2-32B-pref-mix-0325. We use majority voting to improve the quality of answers to our synthetic math questions. For our Persona MATH and Grade School Math… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture-0225.text100K<n<1M22 likes1.1k downloads2y agoHugging Face08PursuitOfDataScience /MiniMax-M2.1-Mixture-of-Thoughts MiniMax-M2.1 Mixture of Thoughts This dataset contains responses generated by MiniMax-M2.1 for user questions from the open-r1/Mixture-of-Thoughts dataset. Dataset Description The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation. Metric Value Examples 349,317 Total Tokens 4,052,592,552 Avg Tokens/Example 11,601 Source Dataset Name:… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts.tabulartext-generation100K<n<1M2 likes956 downloads9mo agoHugging Face09leonli66 /stage3-final-mixturetext10M<n<100M0 likes847 downloads9mo agoHugging Face10akzsh /indic-oss-mixture-cpt-10btext1M<n<10M0 likes824 downloads4mo agoHugging Face11allenai /tulu-3-sft-olmo-2-mixtureNote that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. The OLMo v2 SFT mixture was used to train the OLMo models. It contains 939,344 samples from the following sets: CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024) FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et al., 2023) No Robots (CC-BY-NC-4.0), 9,500… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture.textother100K<n<1M61 likes814 downloads2y agoHugging Face12leonli66 /stage3-final-mixture-cot50 Stage 3 Final Mixture — 50% CoT Compression This is a deterministic capability-preserving rewrite of leonli66/stage3-final-mixture for LCLM Stage-3 post-training. Only the reasoning_data and dolci_think subsets change. Their compression_prompt is the ordinary prompt. A deterministic 50% arm keeps the complete assistant target as ordinary SFT; the other arm wraps the inferred reasoning prefix in <|memory_start|>...<|memory_end|> while keeping the final answer trainable. All… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-final-mixture-cot50.texttext-generation10M<n<100M0 likes800 downloads1mo agoHugging Face13OpenDataArena /ODA-Mixture-500k ODA-Mixture-500k ODA-Mixture-500k is a large-scale general-purpose post-training dataset curated from top-performing open corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination. 🧠 Dataset Summary Domain: General-purpose(e.g., Math, Code, Reasoning, General). Format: Problem → Solution (reasoning trace) → Final answer. Scale (selected training set): ~500K samples. Goal: Achieve maximum general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Mixture-500k.text100K<n<1M123 likes684 downloads5mo agoHugging Face14bguzzo2k /ohlc_1d_mixture Macroeconomic & S&P 500 Yahoo Finance Dataset This repository contains a comprehensive historical dataset for 936 financial instruments, including S&P 500 components, broad market indices, commodities, currencies, and macroeconomic indicators. The data is programmatically extracted from the Yahoo Finance API, cleaned, and normalized for use in quantitative modeling and machine learning. Dataset Hub: bguzzo2k/ohlc_1d_mixture Repository Structure 1d/: Raw daily OHLCV… See the full description on the dataset page: https://huggingface.co/datasets/bguzzo2k/ohlc_1d_mixture.tabular10M<n<100M0 likes666 downloads6mo agoHugging Face15kaizen9 /mixturetext10M<n<100M0 likes620 downloads1y agoHugging Face161231czx /rlhflow_mixture_with_math_del_systemtext1M<n<10M0 likes598 downloads2y agoHugging Face17gokaygokay /lf3-mixture-data lf3-mixture-data Citation and attribution This dataset repository is maintained by Gökay Aydoğan. If you reference this repository in academic work, please cite it as follows and also cite the upstream models, datasets, or projects it builds upon. @dataset{aydogan2026lf3_mixture_data, author = {Aydoğan, Gökay}, title = {{lf3-mixture-data}}, year = {2026}, publisher = {Hugging Face}, url = {https://huggingface.co/datasets/gokaygokay/lf3-mixture-data}… See the full description on the dataset page: https://huggingface.co/datasets/gokaygokay/lf3-mixture-data.image1M<n<10M0 likes466 downloads2mo agoHugging Face18leonli66 /stage3-mixture-cleantext10M<n<100M0 likes376 downloads9mo agoHugging Face19DimitarV /eurospeech-bg-diar-mixtures ⚠️ DEPRECATED — use v2 This dataset contains a shortcut that lets a model infer the number of speakers without listening to the audio. Each speaker was given a fixed 4 turns, so session duration is a direct function of speaker count. Measured on this data: 1-spk median 18.2 s range 11.1-25.2 2-spk 30.4 s 23.0-36.6 3-spk 41.6 s 21.0-54.5 4-spk 54.6 s 38.0-73.0 The 2-speaker and 4-speaker ranges do not overlap — 2-spk tops out… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/eurospeech-bg-diar-mixtures.tabular100K<n<1M0 likes359 downloads1mo agoHugging Face20ZhaoRunyi /Franka2_pour_stir_and_shake_mixture_0330image100K<n<1M0 likes328 downloads6mo agoHugging Face21mxxsc /zip2zip-plus-mixture-partitioned Zip2Zip Plus Mixture Partitioned This dataset is a partitioned pretraining-data mixture built for zip2zip language-model pretraining. The mixture is byte-balanced across four top-level domains: Domain Source Target byte ratio General HuggingFaceFW/fineweb-edu, sample-100BT 50% Code bigcode/the-stack-dedup 20% Math HuggingFaceTB/finemath, finemath-3plus 10% Multilingual epfml/FineWeb2-HQ, 20 language subsets 20% The uploaded layout is partitioned by source… See the full description on the dataset page: https://huggingface.co/datasets/mxxsc/zip2zip-plus-mixture-partitioned.texttext-generation100M<n<1B0 likes326 downloads5mo agoHugging Face22MisterXY89 /SmolLM-lmsys-mixturestext1M<n<10M0 likes325 downloads1y agoHugging Face23dineshananthi /colsmol-mixtureimage100K<n<1M0 likes319 downloads4mo agoHugging Face24OALL /details_jsfs11__MixtureofMerges-MoE-4x7b-v4 Dataset Card for Evaluation run of jsfs11/MixtureofMerges-MoE-4x7b-v4 Dataset automatically created during the evaluation run of model jsfs11/MixtureofMerges-MoE-4x7b-v4. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_jsfs11__MixtureofMerges-MoE-4x7b-v4.tabular100K<n<1M0 likes267 downloads2y agoHugging Face25jdchang /qsharp-full-mixture-1.5b-filtered-with-labelstext10K<n<100K0 likes240 downloads1y agoHugging Face26kamilakesbi /mixture_ami_synthetic_bigaudio1K<n<10K0 likes221 downloads2y agoHugging Face27Thesho1 /ODA-Mixture-500k-copy-58 ODA-Mixture-500k ODA-Mixture-500k is a large-scale general-purpose post-training dataset curated from top-performing open corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination. 🧠 Dataset Summary Domain: General-purpose(e.g., Math, Code, Reasoning, General). Format: Problem → Solution (reasoning trace) → Final answer. Scale (selected training set): ~500K samples. Goal: Achieve maximum general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/Thesho1/ODA-Mixture-500k-copy-58.text100K<n<1M0 likes221 downloads9mo agoHugging Face28OpenRLHF /preference_dataset_mixture2_and_safe_pku Copy from https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku Reward Model Overview This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling . Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0 Model Details If you have any question… See the full description on the dataset page: https://huggingface.co/datasets/OpenRLHF/preference_dataset_mixture2_and_safe_pku.tabular100K<n<1M8 likes210 downloads2y agoHugging Face29OpenDataArena /ODA-Mixture-100k ODA-Mixture-100k ODA-Mixture-100k is a compact general-purpose post-training dataset curated from top-performing open corpora (selected via the *OpenDataArena* leaderboard) and refined through deduplication, benchmark decontamination. 🧠 Dataset Summary Domain: General-purpose(e.g., Math, Code, Reasoning, General). Format: Problem → Solution (reasoning trace) → Final answer. Scale (selected training set): ~100K samples. Goal: Achieve significant general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Mixture-100k.text100K<n<1M99 likes195 downloads8mo agoHugging Face30Thesho1 /ODA-Mixture-500k-copy-53 ODA-Mixture-500k ODA-Mixture-500k is a large-scale general-purpose post-training dataset curated from top-performing open corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination. 🧠 Dataset Summary Domain: General-purpose(e.g., Math, Code, Reasoning, General). Format: Problem → Solution (reasoning trace) → Final answer. Scale (selected training set): ~500K samples. Goal: Achieve maximum general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/Thesho1/ODA-Mixture-500k-copy-53.text100K<n<1M0 likes195 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.