datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
g1_min_episodes_a1_top16_glm47_tracesGLM-4.7-inferredbugs-sandboxes-maxeps-131kg1_selective_top8_diverse_glm47_tracesharbor-devel-sandboxes_glm47_traces_testglm47g1_selective_top8_diverse_100000_glm47_tracesterminal_bench_2_glm46_Toolscale_tasks_traces_20260219_163810g1_min_episodes_top7_85k_glm47_tracesharbor-devel-sandboxes_glm47_sharding_testg1_min_episodes_e1_weighted_top4_100k_glm47_tracesg1_min_episodes_top6_80k_glm47_tracesg1_min_episodes_top8_glm47_tracese1_swegym_combined_capped4_glm47_tracesg1_min_episodes_top8_100k_glm47_tracesg1_weighted_31600_plus_nemotron_10000_glm47_tracese1_gpt_long_swegym_20k_diverse_glm47_tracesGLM4.7-Flash-Nemotronterminal_bench_2_glm46_Toolscale_tasks_traces_20260311_174322e1_gpt_long_swegym_20k_glm47_tracesg1_min_episodes_e1_gpt_long_top8_glm47_tracesg1_min_episodes_d1_all_variants_glm47_tracese1_gpt_long_top50_weighted_top4_glm47_tracesJapanese-Creative-Writing-GLM4.5
Japanese-Creative-Writing-GLM4.5
概要
日本語の小説執筆タスクのデータセットであるAratako/Japanese-Creative-Writing-39.6kから一部の指示を抽出し、zai-org/GLM-4.5で応答を再生成した約8000件のデータセットです。
データセット中の一部データはNSFW表現を含みます。
データの詳細
各データは以下のキーを含みます。
messages: OpenAI messages形式の対話データ
instruction: 指示プロンプト
output: アシスタント応答
system promptは事前に用意した複数種類からランダムに選択されたものが設定されています。
ライセンス
MITライセンスの元配布します。
g1_weighted_31600_cap10_glm47_tracese1_weighted_swesmith_50k_glm47_tracesterminal_bench_2_glm46_Toolscale_tasks_traces_20260310_125924glm-4.7-multiturn-CoT
glm-4.7-multiturn-CoT
Dataset Summary
glm-4.7-multiturn-CoT is a ShareGPT-style multi-turn reasoning distillation dataset generated with GLM-4.7 as the teacher model.
This release focuses on preserving multi-turn dialogue continuity while injecting explicit chain-of-thought style responses in assistant turns.
Key Features
Multi-turn conversation format (human / gpt)
Assistant responses stored as <think>...</think> + final answer
Resume-safe distillation… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/glm-4.7-multiturn-CoT.g1_gptlong_plus_diverse_tezos_glm47_traces
DCAgent2/g1_gptlong_plus_diverse_tezos_glm47_traces
132,259 rows. Concatenation of:
DCAgent/g1_min_episodes_e1_gpt_long_top8_glm47_traces (37,925 rows) — top8 + extra long swegym traces
DCAgent/g1_diverse_tezos_top4_100k_glm47_traces (94,334 rows) — balanced 4-way mix of swesmith/issue/superuser/tezos
No deduplication; shuffled with seed=42. Schema: intersection columns (agent, conversations, date, episode, model, model_provider, result, run_id, task, trial_name).… See the full description on the dataset page: https://huggingface.co/datasets/DCAgent2/g1_gptlong_plus_diverse_tezos_glm47_traces.openthoughts3-1.2m-glm4.7-20kg1_clean_hybrid_scaffold_plus_r2eg_gfi_38k_glm47_traces
