datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
minimax-prompts-samples
MiniMax H3 - 1K
I generated a dataset to test the knowledge scope and capabilities of MiniMax H3.
Samples are of various aspect sizes, and cover a wide range of media types and themes.
The videos are 768 base resolution (~0.6 MP).
They were generated with minimax_h3_fl2va_pruned_int8_convrot.safetensors with 30 steps.
A full video of this dataset can be viewed here https://youtu.be/akkwj9d943Y
lord-of-mysteries-fandom-evidence-sft
Lord of Mysteries Fandom Evidence SFT Dataset
Overview
This dataset provides evidence-aware training and retrieval material for building a Chinese Lord of the Mysteries knowledge assistant.
The release is built from 425 Lord of the Mysteries Fandom Wiki pages. Source URLs, page titles, revision identifiers, and attribution metadata are preserved where available. The companion inference script can retrieve relevant source pages and attach exact source URLs before… See the full description on the dataset page: https://huggingface.co/datasets/xile42/lord-of-mysteries-fandom-evidence-sft.medical_cord19
Description
This dataset contains large amounts of biomedical abstracts and corresponding summaries.
MysteryZebra
Mystery Zebra Dataset
This is the Mystery Zebra dataset created as part of the paper "Lexical Recall or Logical Reasoning: Probing the Limits of Reasoning Abilities in Large Language Models". We make the dataset available in a .csv format for your convenience. The code used to generate the puzzles in this dataset can be found in: https://github.com/arg-tech/MysteryZebra The structure of the dataset is straightforward and easy to parse. In the following, we detail the content of all… See the full description on the dataset page: https://huggingface.co/datasets/arg-tech/MysteryZebra.mystery-agent-cases
Mystery Investigation Cases
Gym-ready investigation case bank for training and evaluating tool-using agents.
Each row is a sealed mystery world: public brief (no solution leak), action catalog with an unlock DAG, evidence items, distractors, and a gold solution. Agents must discover required evidence, then close with a correct culprit and faithful citations — not guess from the story text alone.
Dataset id (example): VaidikML0508/mystery-agent-cases(Use your own HF_DATASET_REPO… See the full description on the dataset page: https://huggingface.co/datasets/VaidikML0508/mystery-agent-cases.nemotron-math-v2-truly-hard-notool
Nemotron-Math-v2 Truly Hard No-Tool Subset
Problems where ALL 6 reasoning regimes score <= 3/8 (37.5%). These are the hardest problems in Nemotron-Math-v2.
Splits
Split
Accuracy
Before (dedup)
After (truly hard)
pass0of8
0/8
6,944
3,217
pass1of8
1/8
3,637
1,321
pass2of8
2/8
4,719
733
pass3of8
3/8
8,363
1,444
Total
23,663
6,715
Filter
All 6 regimes must have accuracy <= 0.375:
reason_high_with_tool / reason_high_no_tool… See the full description on the dataset page: https://huggingface.co/datasets/akh-mysterio/nemotron-math-v2-truly-hard-notool.mystery-crime-booksdeepseek-r1-qwen-32b-planning-mystery-16k
Dataset Card for deepseek-r1-qwen-32b-planning-mystery-16k
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-mystery-16k/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-mystery-16k.nemotron-math-v2-truly-hard-notool-correct-partialdeepseek-r1-qwen-32b-planning-mystery
Dataset Card for deepseek-r1-qwen-32b-planning-mystery
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-mystery/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-mystery.nemotron-math-v2-hard-notool
Nemotron-Math-v2 Hard No-Tool Subset
Filtered from nvidia/Nemotron-Math-v2.
Splits
Split
Rows
Unique Problems
0_8
14,577
6,944
1_8
9,802
3,637
2_8
26,963
4,719
3_8
52,762
8,363
Total
104,104
23,659
employee-burnout-turnover-prediction-800k
Synthetic Employee Dataset
800,000+ employee records with real-world distributions for burnout prediction, turnover analysis, and HR analytics
What You Get
This isn't just another CSV dump. You're looking at 800K+ carefully engineered employee profilesthat mirror actual workforce dynamics, complete with performance metrics, burnout indicators, skill matrices, and behavioral personas. Think of it as a production-ready HR database that never existed but feels like it… See the full description on the dataset page: https://huggingface.co/datasets/Mystic777/employee-burnout-turnover-prediction-800k.qwq-32b-planning-mystery-7-24k
Dataset Card for qwq-32b-planning-mystery-7-24k
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-7-24k/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-7-24k.qwq-32b-planning-mystery-7-24k-greedy
Dataset Card for qwq-32b-planning-mystery-7-24k-greedy
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-7-24k-greedy/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-7-24k-greedy.qwq-32b-planning-mystery-5-24k-greedy
Dataset Card for qwq-32b-planning-mystery-5-24k-greedy
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-5-24k-greedy/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-5-24k-greedy.qwen2.5-32b-instruct-cot-planning-mystery-1-16k-greedy
Dataset Card for qwen2.5-32b-instruct-cot-planning-mystery-1-16k-greedy
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwen2.5-32b-instruct-cot-planning-mystery-1-16k-greedy/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwen2.5-32b-instruct-cot-planning-mystery-1-16k-greedy.qwen2.5-32b-instruct-cot-planning-mystery-4-16k-greedy
Dataset Card for qwen2.5-32b-instruct-cot-planning-mystery-4-16k-greedy
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwen2.5-32b-instruct-cot-planning-mystery-4-16k-greedy/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwen2.5-32b-instruct-cot-planning-mystery-4-16k-greedy.qwq-32b-planning-mystery-4-24k
Dataset Card for qwq-32b-planning-mystery-4-24k
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-4-24k/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-4-24k.qwq-32b-planning-mystery-24k
Dataset Card for qwq-32b-planning-mystery-24k
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-24k/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-24k.llama-3.3-70b-cot-planning-mystery-2-16k-greedy
Dataset Card for llama-3.3-70b-cot-planning-mystery-2-16k-greedy
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/llama-3.3-70b-cot-planning-mystery-2-16k-greedy/raw/main/pipeline.yaml"
or explore the configuration:
distilabel… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/llama-3.3-70b-cot-planning-mystery-2-16k-greedy.deepseek-r1-distill-llama-70b-planning-mystery-1-24k-greedy
Dataset Card for deepseek-r1-distill-llama-70b-planning-mystery-1-24k-greedy
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/deepseek-r1-distill-llama-70b-planning-mystery-1-24k-greedy/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/deepseek-r1-distill-llama-70b-planning-mystery-1-24k-greedy.qwq-32b-planning-mystery-16-24k-greedy
Dataset Card for qwq-32b-planning-mystery-16-24k-greedy
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-16-24k-greedy/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-16-24k-greedy.Mystery_Arena_Results
MysteryArena Results
This dataset stores public MysteryArena run summaries, match records, and compressed episode trajectories for the read-only Streamlit frontend.
The frontend reads index/runs.json first, then loads the unified matches/all_matches.jsonl.gz split. Per-run summaries and compressed trajectory files are kept for metadata and replay.
identifiersllama-3.3-70b-cot-planning-mystery-1-16k-greedy
Dataset Card for llama-3.3-70b-cot-planning-mystery-1-16k-greedy
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/llama-3.3-70b-cot-planning-mystery-1-16k-greedy/raw/main/pipeline.yaml"
or explore the configuration:
distilabel… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/llama-3.3-70b-cot-planning-mystery-1-16k-greedy.qwq-32b-planning-mystery-17-24k-greedy
Dataset Card for qwq-32b-planning-mystery-17-24k-greedy
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-17-24k-greedy/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-mystery-17-24k-greedy.math-hard-traces
Dataset Card for "math-hard-traces"
More Information needed
MysteryWriter
GENERATED DATA SET
This "synthetic" data set was created with the following end user in mind: mystery and crime writers who are working on their next book. This was examined from 4 different perspectives. The data consists of 6,126 questions and answer sets. The tone and approach was set using the following prompt:
Your goal is to help writers with their work, whether they are new or experienced. Word all questions in plain English and maintain a conversational tone. The depth of… See the full description on the dataset page: https://huggingface.co/datasets/theprint/MysteryWriter.llama-3_3-nemotron-super-49b-v1-planning-mystery-15-24k-greedyrust-mystery
