models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-Math-ckpt400DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-Math-DAPO-ckpt180DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-Math-ckpt600DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-MATH-Olympiad-ckpt1200DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-Math-DAPO-ckpt210DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-Math-DAPO-ckpt240DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-Olympiad-Merged-Chains-ckpt1200DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-Math-ckpt150DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-Math-ckpt300DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-Math-ckpt500DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-MATH-Olympiad-ckpt800DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-MATH-Merged-Chains-ckpt300DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-MATH-Merged-Chains-ckpt700DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-MATH-Olympiad-Merged-Chains-ckpt1600DeepSeek-R1-Distill-Qwen-7B-Trajectory-Data-MATH-Olympiad-Merged-Chains-with-prompt-ckpt1200eb_man_sft_trajectory_dataset_baseline_no_reasoning_2epochs-lr1e-5-full-e1-bs-16billiard_trajectory_dataset.jsontrajectory_data.jsontrajectory_data.jsoqwen3-4b-agent-trajectory-loraALFWorld_data_combinationllm-lecture-2025_advanced_make_data_1_model_qwen3-4b-agent-trajectory-lorallm-lecture-2025_advanced_make_data_2_model_qwen3-4b-agent-trajectory-lorallm-lecture-2025_advanced_make_data_3_model_qwen3-4b-agent-trajectory-loraqwen3-4b-agent-trajectory-DataMerge_rev.7_rev.6_lora-SFT-SQL-ALFWorld_rev.0.6robot-trajectory-dataset
