datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GLM-4.7-358B-logits
Dataset Description
This dataset contains model logits extracted from zai-org/GLM-4.7-FP8 using DistillKit.
Source Data
Logits were generated from prompts drawn from the following publicly available datasets:
BEE-spoke-data/fineweb-100k_en-med
flytech/python-codes-25k
agentlans/multilingual-text
These sources collectively provide a mix of English web text, programming-related content, and multilingual natural language data.
Intended Use
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/JackBinary/GLM-4.7-358B-logits.glm46v-flash-ultramega-sparse-hmapg1_min_episodes_a1_top16_glm47_tracescc_glm4.7_distillGLM-4.7-inferredbugs-sandboxes-maxeps-131kg1_selective_top8_diverse_glm47_tracesharbor-devel-sandboxes_glm47_traces_testglm47g1_selective_top8_diverse_100000_glm47_tracesterminal_bench_2_glm46_Toolscale_tasks_traces_20260219_163810g1_min_episodes_top7_85k_glm47_tracesharbor-devel-sandboxes_glm47_sharding_testg1_min_episodes_e1_weighted_top4_100k_glm47_tracesg1_min_episodes_top6_80k_glm47_tracesg1_min_episodes_top8_glm47_tracese1_swegym_combined_capped4_glm47_tracesg1_min_episodes_top8_100k_glm47_tracesg1_weighted_31600_plus_nemotron_10000_glm47_tracese1_gpt_long_swegym_20k_diverse_glm47_tracesGLM4.7-Flash-Nemotronterminal_bench_2_glm46_Toolscale_tasks_traces_20260311_174322e1_gpt_long_swegym_20k_glm47_tracesg1_min_episodes_e1_gpt_long_top8_glm47_tracesg1_min_episodes_d1_all_variants_glm47_tracese1_gpt_long_top50_weighted_top4_glm47_tracesJapanese-Creative-Writing-GLM4.5
Japanese-Creative-Writing-GLM4.5
概要
日本語の小説執筆タスクのデータセットであるAratako/Japanese-Creative-Writing-39.6kから一部の指示を抽出し、zai-org/GLM-4.5で応答を再生成した約8000件のデータセットです。
データセット中の一部データはNSFW表現を含みます。
データの詳細
各データは以下のキーを含みます。
messages: OpenAI messages形式の対話データ
instruction: 指示プロンプト
output: アシスタント応答
system promptは事前に用意した複数種類からランダムに選択されたものが設定されています。
ライセンス
MITライセンスの元配布します。
g1_weighted_31600_cap10_glm47_tracese1_weighted_swesmith_50k_glm47_tracesterminal_bench_2_glm46_Toolscale_tasks_traces_20260310_125924glm-4.7-multiturn-CoT
glm-4.7-multiturn-CoT
Dataset Summary
glm-4.7-multiturn-CoT is a ShareGPT-style multi-turn reasoning distillation dataset generated with GLM-4.7 as the teacher model.
This release focuses on preserving multi-turn dialogue continuity while injecting explicit chain-of-thought style responses in assistant turns.
Key Features
Multi-turn conversation format (human / gpt)
Assistant responses stored as <think>...</think> + final answer
Resume-safe distillation… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/glm-4.7-multiturn-CoT.
