datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glm-4.7-multiturn-CoT
glm-4.7-multiturn-CoT
Dataset Summary
glm-4.7-multiturn-CoT is a ShareGPT-style multi-turn reasoning distillation dataset generated with GLM-4.7 as the teacher model.
This release focuses on preserving multi-turn dialogue continuity while injecting explicit chain-of-thought style responses in assistant turns.
Key Features
Multi-turn conversation format (human / gpt)
Assistant responses stored as <think>...</think> + final answer
Resume-safe distillation… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/glm-4.7-multiturn-CoT.GLM-4-Instruct-4K-zh
Dataset Card for Dataset Name
❤️欢迎使用rqq/GLM-4-Instruct-4K-zh数据集,本数据集包含了4000条高质量的glm4回复。
该数据集的提问数据源自高质量的Sao10K/Claude-3-Opus-Instruct-5K数据集,我们把它的问题翻译成了中文,使用glm-4进行了重新回答。
该数据集使用alpaca格式,可以直接用在llama-factory项目中进行训练!
文件如下:
GLM-4-Instruct-4K-zh.json 问答数据集,alpaca格式
GLM-4-question-translate-5K-zh 翻译-对话数据集,记录了把Sao10K/Claude-3-Opus-Instruct-5K问题翻译成中文的数据
Welcome to the rqq/GLM-4-Instruct-4K-zh dataset! This dataset includes 4,000 high-quality responses from the GLM-4 model.
The question data… See the full description on the dataset page: https://huggingface.co/datasets/rqq/GLM-4-Instruct-4K-zh.glm-4.7-Superior-Reasoning-stage1
glm-4.7-Superior-Reasoning-stage1
Dataset Summary
glm-4.7-Superior-Reasoning-stage1 is a Stage1 reasoning distillation dataset built from the Alibaba Superior-Reasoning style pipeline, with a stronger teacher model replacement.
Compared with the original upstream setup, this release uses GLM-4.7 as teacher for higher-quality reasoning traces.
Stage1 Distillation Setup (Low Temperature)
Training stage: stage1
Sampling temperature: 0.6 (low-temperature… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/glm-4.7-Superior-Reasoning-stage1.
