llm-reward
details_Lichang-Chen__reward_max_spin_filter0.7logic_llm_reward_may25llm-reward-gen-quad-sft-7095llm-reward-generator-huyen-quadcopter-sft-889llm-reward-generator-huyen-quadcopter-dpo-111llm-reward-generator-quadcopter-sft-v3-huyen-889
LLM Reward Generator Quadcopter SFT V3 Huyen 889 (Updated over V2)
Removed system messages
All dataset rows now contain only user and assistant messages. The system role has been removed from both SFT and DPO outputs.
Removed few-shot example from prompt
The <valid_example> code block that was previously included in the user prompt has been removed. The model now receives only the task description and tensor reference — no reference implementation… See the full description on the dataset page: https://huggingface.co/datasets/UPB-RAT-Lab/llm-reward-generator-quadcopter-sft-v3-huyen-889.
