glm-4.5
Datasets
All datasets matching “glm-4.5”Medical-Reasoning-SFT-GLM_4.5_Air
Medical-Reasoning-SFT-GLM_4.5_Air
A large-scale medical reasoning dataset generated using zai-org/GLM-4.5-Air, containing over 225,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
zai-org/GLM-4.5-Air
Total Samples
225,179
Samples with Reasoning
224,942 (99.9%)
Estimated Tokens
~441 Million
Content Tokens
~315 Million
Reasoning Tokens
~126 Million
Language
English… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GLM_4.5_Air.Japanese-Creative-Writing-GLM4.5
Japanese-Creative-Writing-GLM4.5
概要
日本語の小説執筆タスクのデータセットであるAratako/Japanese-Creative-Writing-39.6kから一部の指示を抽出し、zai-org/GLM-4.5で応答を再生成した約8000件のデータセットです。
データセット中の一部データはNSFW表現を含みます。
データの詳細
各データは以下のキーを含みます。
messages: OpenAI messages形式の対話データ
instruction: 指示プロンプト
output: アシスタント応答
system promptは事前に用意した複数種類からランダムに選択されたものが設定されています。
ライセンス
MITライセンスの元配布します。
Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High
Distill
This is a multi-source curated instruction and reasoning dataset specifically for training and distilling large language models (LLMs) to exhibit advanced Chain-of-Thought (CoT), Agentic, Mathematical and Coding capabilities. It aggregates high-quality outputs from frontier models into messages ChatML format.
Dataset Structure
The dataset contains a total of 70.2K examples, split into three subsets based on the presence of visible reasoning… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High.arxiv-sample-affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air-inference-results-enriched
affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air arXiv author affiliation inference results
Author names and institutional affiliations extracted from arXiv preprints with the affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air LoRA, enriched with ROR identifiers.
Dataset Structure
Each record contains the following fields:
Field
Type
Description
doi
string
DOI for the preprint
title
string
Preprint title
arxiv_id
stringarXiv identifier… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-sample-affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air-inference-results-enriched.when2call-GLM4.5-IItau2-bench_zai-org_GLM-4.5-Air_n1_r1
