zuxler/qwen2.5-1.5b-legal-grpo-reasoning
0531
(Trained with Unsloth)
(Trained with Unsloth)
(Trained with Unsloth)
Unsloth Model Card
(Trained with Unsloth)
(Trained with Unsloth)
(Trained with Unsloth)
(Trained with Unsloth)
Unsloth Model Card
initial commit
