datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lord-of-mysteries-fandom-evidence-sft
Lord of Mysteries Fandom Evidence SFT Dataset
Overview
This dataset provides evidence-aware training and retrieval material for building a Chinese Lord of the Mysteries knowledge assistant.
The release is built from 425 Lord of the Mysteries Fandom Wiki pages. Source URLs, page titles, revision identifiers, and attribution metadata are preserved where available. The companion inference script can retrieve relevant source pages and attach exact source URLs before… See the full description on the dataset page: https://huggingface.co/datasets/xile42/lord-of-mysteries-fandom-evidence-sft.mystery-agent-cases
Mystery Investigation Cases
Gym-ready investigation case bank for training and evaluating tool-using agents.
Each row is a sealed mystery world: public brief (no solution leak), action catalog with an unlock DAG, evidence items, distractors, and a gold solution. Agents must discover required evidence, then close with a correct culprit and faithful citations — not guess from the story text alone.
Dataset id (example): VaidikML0508/mystery-agent-cases(Use your own HF_DATASET_REPO… See the full description on the dataset page: https://huggingface.co/datasets/VaidikML0508/mystery-agent-cases.the_pile_mysticThe Pile is a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality
datasets combined together.math
Mystery Machine - Math Reasoning Data
Training and evaluation data for the math expert of the CS-552 (Spring 2026) Mystery Machine team's Qwen3-1.7B post-training project. This dataset bundles every data artifact used to produce and evaluate the shipped math model: the raw RL problem pool, the two learnability-filtered training sets corresponding to the two RL phases we ran, the supervised fine-tuning corpus (including chain-of-thought traces we distilled ourselves from a larger… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-mystery-machine/math.MYSTLYE
Erik Merkel
MyStarData-Medium-5M-LLaMA-3.1-Ins-Tokenizemy-stem-index
