sambal/k2-think-v2-grpo-checkpoint-320
0468
K2-Think-V2 GRPO — checkpoint 320
Merged Hugging Face checkpoint from global step 320 of K2-Think-V2 GRPO training.
- Architecture:
LlamaForCausalLM - Weight precision: FP32 (
float32), preserved from the merged checkpoint - Format: full model weights, split across 62 safetensors shards (not a LoRA adapter)
- Includes the checkpoint's original configuration, tokenizer, and chat template
The checkpoint is uploaded without quantization or changes to its weights or tokenizer.
