datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prompt-swap-mixed12-5xlr-e1-mxfp4-mergedprompt-swap-mixed12-5xlr-e2-mxfp4-mergedMiniMax-M3-150k-Mixed
m3-alldomains-verified-107k
Verified distillation traces generated with faststill v0.0.1 — a pipeline that generates (prompt, reasoning, output) triplets from any OpenAI-compatible chat-completions endpoint and deterministically verifies every row before keeping it. A row is verified=true only when a machine check (executed unit tests, exact / normalized answer compare) confirmed it, so wrong labels are filtered out instead of poisoning a student model.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/MiniMax-M3-150k-Mixed.mixed-pretrain-10b-gpt2
Mixed Pretraining 10B (GPT-2 BPE)
A 10-billion-token pretraining dataset, GPT-2 BPE tokenized, assembled as a
diverse mix of web text, books, Wikipedia, code, academic papers, Q&A and
instruction-formatted conversations.
Built to train a ~500M parameter from-scratch GPT-2-style transformer (see
juliannunezb/transformer-lm-500m).
Mix
Source
Mix %
Tokens
Notes
fineweb
40.4%
4,039,999,700
reused from kjj0/fineweb10B-gpt2
fineweb_edu
15.2%
1,514,999,900
reused… See the full description on the dataset page: https://huggingface.co/datasets/juliannunezb/mixed-pretrain-10b-gpt2.minimax-m3-150k-mixed
m3-alldomains-verified-107k
Verified distillation traces generated with faststill v0.0.1 — a pipeline that generates (prompt, reasoning, output) triplets from any OpenAI-compatible chat-completions endpoint and deterministically verifies every row before keeping it. A row is verified=true only when a machine check (executed unit tests, exact / normalized answer compare) confirmed it, so wrong labels are filtered out instead of poisoning a student model.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/minimax-m3-150k-mixed.Samvaad-Mixed-Language-2prompt-difficulty-mixed
Prompt Difficulty Meta-Analysis
Introduction
The difficulty of large language model (LLM) prompts varies widely, from simple queries to complex multi-step reasoning tasks.
This study develops a consistent, data-driven difficulty score for English ChatGPT prompts, using classifiers trained on labelled difficulty datasets.
The goal is to improve automated prompt difficulty classification.
Methods
Detailed methods
Several methods were used to quantify the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-difficulty-mixed.AXXXX_jssp_mixed_step_train_dispatch_v1Samvaad-Mixed-Language-3laion2b-mixed-1024px-human
Overview
This dataset is a selective merge of some other of our datasets. Mainly, I pulled human-centric, real-world photos from the following datasets:
opendiffusionai/laion2b-45ish-1120px
opendiffusionai/laion2b-squareish-1024px
opendiffusionai/laion2b-23ish-1216px
As such, it is "mixed" aspect ratio.
The very smallest height ones are from our "squarish" set, so are at least 1024px tall.
However, the other ones with longer rations, have an appropriately longer minimum… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-mixed-1024px-human.S3-CoT-Self-Sampled-Data-v2_mixed
Dataset Overview
We release a new self-sampled variable-length reasoning dataset based on DeepSeek-R1-Distill-Qwen-7B.
The seed problems are collected from three high-quality datastes: MATH500, LIMO-v2, and AIME problems released before 2023.
The data are used for efficient cot learning in our paper (S3-CoT: Self-Sampled Succinct Reasoning Enables Efficient Chain-of-Thought LLMs).
For each instruction (problem), we provide multiple reasoning traces generated under different… See the full description on the dataset page: https://huggingface.co/datasets/yrdu/S3-CoT-Self-Sampled-Data-v2_mixed.mixed_data_onesokoban_adaptive_3box_hard_mixed_350
sokoban_adaptive_3box_hard_mixed_350
BAGEL VLM-Gym world-model dataset (sokoban / eval).
350 hard 3-box episodes: deadlock + trivial mixed in one contiguous pool; train-disjoint adaptive-thinking (when-to-think) benchmark.
layout: Held-out eval episode pool. shard_*.jsonl.gz at the repo root; one episode per row with a contiguous global index.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching sokoban checkpoint(s) under the… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/sokoban_adaptive_3box_hard_mixed_350.sokoban_easy_mixed_deadlock_trivial
sokoban_easy_mixed_deadlock_trivial
BAGEL VLM-Gym world-model dataset (sokoban / eval).
Deadlock (temptation tiers) + trivial (no-deadlock) episodes in one contiguous pool; train-disjoint elicitation eval.
layout: Held-out eval episode pool. shard_*.jsonl.gz at the repo root; one episode per row with a contiguous global index.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching sokoban checkpoint(s) under the companion model org; CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/sokoban_easy_mixed_deadlock_trivial.mvdsc_mixed_for_ml@inproceedings{10.1145/3488932.3527288,
author = {Zhou, Xin and Verma, Rakesh M.},
title = {Vulnerability Detection via Multimodal Learning: Datasets and Analysis},
year = {2022},
isbn = {9781450391405},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3488932.3527288},
doi = {10.1145/3488932.3527288},
abstract = {A vulnerability is a weakness that can be exploited by an attacker, e.g., performing unauthorized actions within a… See the full description on the dataset page: https://huggingface.co/datasets/ijakenorton/mvdsc_mixed_for_ml.Samvaad-Mixed-Languagemixed_data_zeroAXXXX_jssp_mixed_step_train_all_v1data_fix_before_jssp_mixed_step_train_dispatch_v1data_fix_before_jssp_mixed_step_train_all_v1
