datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-SWE-smith-Rust-GLM-4.6-trajectoriesharbor-devel-sandboxes_glm_4.6_traces_openhandscombined-reasoning-opus-4.6-opus-4.7-kimi-k2.5-kimi-k2.6-glm-5.1
Combined Reasoning Distill — Multi-Model
A large-scale unified reasoning dataset combining thinking and chain-of-thought traces distilled from frontier models, normalized into a single consistent schema for fine-tuning. Includes data from Claude (Opus 4.5/4.6/4.7, Sonnet 4.5/4.6, Haiku 4.5), GPT (5.1/5.2), Gemini 3 Pro Preview, Kimi (K2/K2.5/K2.6), GLM (4.6/4.7/5.1), MiniMax M2.1, Grok Code Fast 1, and more.
Schema
Every row has a single field:
Field
Type… See the full description on the dataset page: https://huggingface.co/datasets/Avtrkrb/combined-reasoning-opus-4.6-opus-4.7-kimi-k2.5-kimi-k2.6-glm-5.1.DCAgent2_terminal_bench_2_DCAgent2_nl2bash-verified-GLM-4.6-traces-32ep-32k_glof5f5a963DCAgent2_terminal_bench_2_DCAgent2_nl2bash-verified-GLM-4.6-traces-32ep-32k_glob0c37db0DCAgent2_terminal_bench_2_DCAgent2_nl2bash-verified-GLM-4.6-traces-32ep-32k_glocbb0258aDCAgent2_terminal_bench_2_DCAgent_nl2bash-GLM-4.6-traces_Qwen3-8B_20251120_153822Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High
Distill
This is a multi-source curated instruction and reasoning dataset specifically for training and distilling large language models (LLMs) to exhibit advanced Chain-of-Thought (CoT), Agentic, Mathematical and Coding capabilities. It aggregates high-quality outputs from frontier models into messages ChatML format.
Dataset Structure
The dataset contains a total of 70.2K examples, split into three subsets based on the presence of visible reasoning… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High.SWE-smith-rs-glm-4.6-trajectoriesDCAgent2_terminal_bench_2_DCAgent2_nl2bash-verified-GLM-4.6-traces-32ep-32k-ab1674537b6DCAgent2_terminal_bench_2_DCAgent2_nl2bash-verified-GLM-4.6-traces-32ep-32k-ab13d395731DCAgent2_terminal_bench_2_DCAgent2_nl2bash-verified-GLM-4.6-traces-32ep-32k-ab18e4c9bdfDCAgent2_terminal_bench_2_DCAgent2_nl2bash-verified-GLM-4.6-traces-32ep-32k-ab116928ab6DCAgent2_terminal_bench_2_DCAgent_code_contests-GLM-4.6-traces_Qwen3-8B_20251124_180143DCAgent2_terminal_bench_2_DCAgent_code_contests-GLM-4.6-traces_Qwen3-8B_20251201_113400DCAgent2_swebench-verified-random-100-folders_DCAgent_nl2bash-GLM-4.6-traces_Qw376a20a1DCAgent2_swebench-verified-random-100-folders_DCAgent_nl2bash-GLM-4.6-traces_Qw3a3e0403DCAgent2_swebench-verified-random-100-folders_DCAgent2_nl2bash-verified-GLM-4.6fbdaa14eDCAgent2_swebench-verified-random-100-folders_DCAgent2_nl2bash-verified-GLM-4.67a4c5141swesmith-GLM-4.6-32ep-32k-v2-traces-chunk002DCAgent2_swebench-verified-random-100-folders_DCAgent2_nl2bash-verified-GLM-4.6ef4a71a5DCAgent2_terminal_bench_2_laion_nl2bash-verified-GLM-4.6-traces-32ep-32k-mgn1e37bf05c35nl2bash-GLM-4.6-tracesDCAgent2_terminal_bench_2_laion_nl2bash-verified-GLM-4.6-traces-32ep-32k-mgn1e552b75b5astackexchange-tezos-sandboxes_glm_4.6_traces_locetashDCAgent2_terminal_bench_2_laion_nl2bash-verified-GLM-4.6-traces-32ep-32k-mgn5e492059f8fDCAgent2_terminal_bench_2_laion_GLM-4.6-stackoverflow-32eps-65k-fixeps_Qwen3-8B2074e6c1stackexchange-tezos-sandboxes_glm_4.6_traces_locetash_againswesmith-GLM-4.6-32ep-32k-v2-traces-chunk000DCAgent2_swebench-verified-random-100-folders_DCAgent2_nl2bash-verified-GLM-4.674c92fa7
