ansulev/glm-5.1-reasoning-1m-cleaned
GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data:… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/glm-5.1-reasoning-1m-cleaned.
GLM-5.1-Reasoning-1M-Cleaned

GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary

- Teacher model in the data: GLM-5.1
- Total processed records: 766,535
- Total kept records: 746,321
- Total removed records: 20,214
- Subset layout preserved exactly as the source dataset:
main,PHD-Science,Multilingual-STEM,Math
Included Content
main: general reasoning and instruction-following data.PHD-Science: graduate-level physics, chemistry, and biology reasoning traces.Multilingual-STEM: multilingual STEM reasoning data, including Chinese, English, and other languages present in the source release.Math: mathematics-heavy reasoning and proof-style responses.
Cleaning and Reformatting
The raw source dataset mixed two answer layouts:
- Standard
<think>...</think>reasoning tags. - A non-standard short-dash wrapper around the reasoning section.
This cleaned release normalizes both styles into a single output format:
{
"id": "md5-hash-of-domain-input-reasoning-answer",
"conversations": [
{"from": "human", "value": "user prompt"},
{"from": "gpt", "value": "<think>\nreasoning trace\n</think>\n\nfinal answer"}
],
"input": "user prompt",
"output": "<think>\nreasoning trace\n</think>\n\nfinal answer",
"domain": "subset/domain name from the original _id prefix",
"meta": {
"input_tokens": 123,
"output_tokens": 456,
"teacher_model": "GLM-5.1"
}
}Removed data
The cleaning pipeline removed records with:
- incomplete or obviously truncated answers,
- repeated reasoning paragraphs or duplicated answer segments,
- refusal-style answers,
- unparseable reasoning/answer boundaries,
- exact duplicate records after normalization.
Subset Statistics
Filter Statistics
Additional Token Statistics
Data Structure
Each example is a single-turn reasoning distillation sample:
conversations[0]: the user prompt.conversations[1]: the model response with reasoning wrapped in<think>...</think>and the final answer after the closing tag.input: prompt-only view for training pipelines that prefer flat prompt fields.output: tagged answer-only view for training pipelines that prefer flat completion fields.domain: original subset/domain name extracted from the source record ID.meta: lightweight per-example metadata.
Usage
from datasets import load_dataset
main = load_dataset("Jackrong/GLM5.1-Reasoning-1M-Cleaned", "main")
science = load_dataset("Jackrong/GLM5.1-Reasoning-1M-Cleaned", "PHD-Science")
stem = load_dataset("Jackrong/GLM5.1-Reasoning-1M-Cleaned", "Multilingual-STEM")
math = load_dataset("Jackrong/GLM5.1-Reasoning-1M-Cleaned", "Math")If you publish under a different namespace, replace Jackrong/ with your actual Hugging Face username or org.
Provenance
This dataset is derived from:
- Original dataset: `Kassadin88/GLM-5.1-1000000x`
- Original author: Kassadin88
Citation
Please cite the original dataset first:
@misc{glm51-1000000x,
title={GLM-5.1-1000000x: One Million Reasoning Traces Distilled from GLM-5.1},
author={Kassadin88},
year={2026},
publisher={HuggingFace},
url={https://huggingface.co/datasets/Kassadin88/GLM-5.1-1000000x}
}You can additionally cite this cleaned derivative release as:
@misc{glm51_reasoning_1m_cleaned,
title={GLM-5.1-Reasoning-1M-Cleaned},
author={Jackrong},
year={2026},
publisher={HuggingFace},
url={https://huggingface.co/datasets/Jackrong/GLM5.1-Reasoning-1M-Cleaned}
}