ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO-dataset_for_DPO
OminiGAIA-DPO-data This dataset contains the final DPO training pairs used to train ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO. The pairs are derived from 7B-native rollouts on OmniGAIA train questions: Roll out the SFT model on answer-hidden train inputs. Audit each rollout with Gemini using the private reference answer and annotated solution. Locate the first erroneous assistant sub-step. Convert the corrected prefix (tau_win) and the original erroneous prefix… See the full description on the dataset page: https://huggingface.co/datasets/ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO-dataset_for_DPO.
OminiGAIA-DPO-data
This dataset contains the final DPO training pairs used to train ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO.
The pairs are derived from 7B-native rollouts on OmniGAIA train questions:
- Roll out the SFT model on answer-hidden train inputs.
- Audit each rollout with Gemini using the private reference answer and annotated solution.
- Locate the first erroneous assistant sub-step.
- Convert the corrected prefix (
tau_win) and the original erroneous prefix (tau_lose) into ms-swift DPO format.
Files
Dataset statistics
Conversion skip reasons:
{
"ok": 962,
"trace_is_correct": 821,
"parse_failed": 25,
"outer_idx_out_of_range": 71,
"target_not_assistant": 3
}Format
Each JSONL row is compatible with ms-swift RLHF/DPO training and includes:
idmessagesrejected_messagesimagesaudiosvideosrejected_imagesrejected_audiosrejected_videosanswerrollout_predicted_answerrollout_statusfirst_error_indexerror_reason
Caveats
- This is derived preference data, not the full raw rollout/audit trace release.
- One rollout shard was excluded because it repeatedly timed out.
- The corrected branch contains Gemini-generated corrected sub-steps; tool observations were not re-executed and inserted into the chosen branch.
- OmniGAIA train/test share a media pool by benchmark design. The upstream SFT checkpoint verifies no test-set question text appears verbatim in training samples.
Intended use
Use this dataset to reproduce or inspect the DPO stage of ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO.
It is not intended as a general-purpose RLHF dataset outside OmniGAIA-style omni-modal tool-use agents.
