nabin2004/AOS-qwen3-8b-narrated-dpo
0
AOS-qwen3-8b-narrated-dpo
Direct Preference Optimization (DPO) aligned LoRA adapter for Manim Community Edition mathematical and educational animation synthesis.
Alignment Objective
Aligns the model to strongly prefer generating synchronized `manim-voiceover` code examples:
- Chosen: Clean
VoiceoverScenescripts withself.set_speech_service(GTTSService()), animation duration tracking (run_time=tracker.duration), and phonetic mathematical explanations. - Rejected: Silent, un-narrated standard
Scenecode.
Lineage
- Base LLM:
Qwen/Qwen3-8B - SFT Prior:
nabin2004/AOS-qwen3-8b-narrated-adapter
