CoolFace
Modelpublic

sambal/k2-think-v2-grpo-checkpoint-320

sourceHugging Faceupdated 8d agoView on Hugging Face
0likes468downloads
Model Card

K2-Think-V2 GRPO — checkpoint 320

Merged Hugging Face checkpoint from global step 320 of K2-Think-V2 GRPO training.

  • —Architecture: LlamaForCausalLM
  • —Weight precision: FP32 (float32), preserved from the merged checkpoint
  • —Format: full model weights, split across 62 safetensors shards (not a LoRA adapter)
  • —Includes the checkpoint's original configuration, tokenizer, and chat template

The checkpoint is uploaded without quantization or changes to its weights or tokenizer.