CoolFace
Modelpublic

helloAK96/chaosops-grpo-lora-p3a

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes11downloads
9 commits on main
a2a33175mo ago

Fix model card frontmatter: base_model -> Qwen/Qwen2.5-3B-Instruct

helloAK96
75959b15mo ago

Phase 3A model card: full eval results, training recipe, trade-off notes

helloAK96
60a270b5mo ago

Add post-training evaluation.json

helloAK96
72f44665mo ago

Add post-training evaluation_summary.txt

helloAK96
d6be9e05mo ago

Add post-training comparison_curve.png

helloAK96
f06302d5mo ago

Add learning_curve.png

helloAK96
ac8afc65mo ago

Add training_metrics.json

helloAK96
05f402f5mo ago

Upload GRPO LoRA adapter

helloAK96
13cb9315mo ago

initial commit

helloAK96