datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glm-4.7-multiturn-CoT
glm-4.7-multiturn-CoT
Dataset Summary
glm-4.7-multiturn-CoT is a ShareGPT-style multi-turn reasoning distillation dataset generated with GLM-4.7 as the teacher model.
This release focuses on preserving multi-turn dialogue continuity while injecting explicit chain-of-thought style responses in assistant turns.
Key Features
Multi-turn conversation format (human / gpt)
Assistant responses stored as <think>...</think> + final answer
Resume-safe distillation… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/glm-4.7-multiturn-CoT.cc_glm47_distillGLM-4-Instruct-4K-zh
Dataset Card for Dataset Name
❤️欢迎使用rqq/GLM-4-Instruct-4K-zh数据集,本数据集包含了4000条高质量的glm4回复。
该数据集的提问数据源自高质量的Sao10K/Claude-3-Opus-Instruct-5K数据集,我们把它的问题翻译成了中文,使用glm-4进行了重新回答。
该数据集使用alpaca格式,可以直接用在llama-factory项目中进行训练!
文件如下:
GLM-4-Instruct-4K-zh.json 问答数据集,alpaca格式
GLM-4-question-translate-5K-zh 翻译-对话数据集,记录了把Sao10K/Claude-3-Opus-Instruct-5K问题翻译成中文的数据
Welcome to the rqq/GLM-4-Instruct-4K-zh dataset! This dataset includes 4,000 high-quality responses from the GLM-4 model.
The question data… See the full description on the dataset page: https://huggingface.co/datasets/rqq/GLM-4-Instruct-4K-zh.glm-4.7-2000xThis is a reasoning dataset created using GLM 4.7 with reasoning effort set to high (not sure if that flag does anything for this model though).
The dataset is meant for creating distilled versions of GLM 4.7 by fine-tuning already existing open-source LLMs.
This dataset includes an addition to our recently enhanced set of prompts to cover creative writing multilingual reasoning and various graduate level questions across a wide variety of domains.
Stats
Costs: $ 30.72 (USD)… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/glm-4.7-2000x.glm-4.7-Superior-Reasoning-stage1
glm-4.7-Superior-Reasoning-stage1
Dataset Summary
glm-4.7-Superior-Reasoning-stage1 is a Stage1 reasoning distillation dataset built from the Alibaba Superior-Reasoning style pipeline, with a stronger teacher model replacement.
Compared with the original upstream setup, this release uses GLM-4.7 as teacher for higher-quality reasoning traces.
Stage1 Distillation Setup (Low Temperature)
Training stage: stage1
Sampling temperature: 0.6 (low-temperature… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/glm-4.7-Superior-Reasoning-stage1.glm47-synth-v1-dataset
GLM47 Synth V1 Dataset
This repository contains 260 verified Aider-format supervised fine-tuning rows. The dataset has ten synthetic variants for each of 26 C++ task families and was structured for an exact 100-epoch memorization experiment.
Dataset structure
The dataset has one configuration (default) and one split (train):
File
Rows
Format
sft/train.jsonl
260
UTF-8 JSON Lines
Every row contains these top-level fields:
label: unique row label… See the full description on the dataset page: https://huggingface.co/datasets/HimanshuPathak/glm47-synth-v1-dataset.glm47-calibration-1360
[!TIP]
Support this work: donate.sybilsolutions.ai
REAP surfaces: GLM | MiniMax | Qwen | Gemma | Paper | Code | PR17 | Cerebras Collection
GLM-4.7 REAP Calibration Dataset (1360 samples)
Mixed calibration dataset for GLM-4.7-REAP-218B W4A16 quantization.
Composition
Code generation: 700 samples (evol-codealpaca-v1 style)
Function calling: 330 samples (xlam-function-calling-60k style)
Agentic trajectories: 330 samples (SWE-smith-trajectories style)… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/glm47-calibration-1360.glm-4.7-multiturn-cot-jackrong-openai-formatglm-4.6-250xThis is a reasoning dataset created using GLM 4.6. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of 4.6 by fine-tuning already existing open-source LLMs.
Disclaimer: With such a small dataset any fine-tuned LLM will only be reproducing the answer/thinking style. No knowledge transfer is happening when fine-tuning on this dataset.
glm-4.7-2048-reasoning-1000xglm-4.7-350x
GLM 4.7 - 350x
This is a reasoning dataset created using GLM 4.7. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of 4.7 by fine-tuning already existing open-source LLMs.
Stats
Cost: $ 1.60 (USD)
Total Tokens: $ 1.10 M
glm47-pie-cpp-posttraining-data
GLM-4.7-Flash PIE C++ Post-Training Data
The exact prepared dataset used for the GLM-4.7-Flash C++ performance
post-training runs.
Splits
File
Rows
Purpose
sft/train.jsonl
7,864
Supervised fine-tuning
grpo/train.jsonl
7,887
GRPO prompt and reward evaluation
eval/validation.jsonl
1,259
Full held-out evaluation
eval/validation_mini126.jsonl
126
Fast evaluation
eval/validation_mini4.jsonl
4
Smoke evaluation
tasks.tar.gz
9,146 task JSONs
Reward… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/glm47-pie-cpp-posttraining-data.best-dataset-glm47flashglm47-synth-v1-dataset
Synth v1 Dataset
This package contains 260 verified Aider-format SFT rows: ten synthetic
variants for each of the 26 source task families. It is intentionally built for
an exact 100-epoch memorization experiment.
The training file is sft/train.jsonl. Every row uses the same nine-message
aider-chat-v1 structure as the successful SFT-v5 package. Tests are not
model-visible; the independent verifier replays each final assistant target
against its source C++ test suite.… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/glm47-synth-v1-dataset.GLM4.6-OpenR1Math-SFTglm47-aider-rl8-validity-rollouts-20260723Glm4.7-sanguoglm4.7glm47-aider-full-v5-rl
