lightshifted/llama-3.1-8B-GRPO-rag-rewards
08
Trained with Unsloth
Trained with Unsloth
Trained with Unsloth
Upload tokenizer
Upload README.md with huggingface_hub
initial commit
Trained with Unsloth
Trained with Unsloth
Trained with Unsloth
Upload tokenizer
Upload README.md with huggingface_hub
initial commit