CoolFace
Datasetpublic

ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO-dataset_for_DPO

OminiGAIA-DPO-data This dataset contains the final DPO training pairs used to train ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO. The pairs are derived from 7B-native rollouts on OmniGAIA train questions: Roll out the SFT model on answer-hidden train inputs. Audit each rollout with Gemini using the private reference answer and annotated solution. Locate the first erroneous assistant sub-step. Convert the corrected prefix (tau_win) and the original erroneous prefix… See the full description on the dataset page: https://huggingface.co/datasets/ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO-dataset_for_DPO.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
1likes24downloads
Dataset Card

OminiGAIA-DPO-data

This dataset contains the final DPO training pairs used to train ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO.

The pairs are derived from 7B-native rollouts on OmniGAIA train questions:

  1. 1.Roll out the SFT model on answer-hidden train inputs.
  2. 2.Audit each rollout with Gemini using the private reference answer and annotated solution.
  3. 3.Locate the first erroneous assistant sub-step.
  4. 4.Convert the corrected prefix (tau_win) and the original erroneous prefix (tau_lose) into ms-swift DPO format.

Files

FileDescription
omnidpo_dpo_train_7B_7shards.jsonlFinal ms-swift DPO dataset (messages / rejected_messages).
omnidpo_dpo_train_7B_7shards.report.jsonConversion report and skip statistics.

Dataset statistics

ItemValue
Source model for rolloutsZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-only
Clean rollouts audited1,882
Completed rollout shards7 / 8
Final DPO pairs962
Pair formatmessages = corrected prefix, rejected_messages = original erroneous prefix
VerifierGemini audit with access to private answer and annotated solution

Conversion skip reasons:

json
{
  "ok": 962,
  "trace_is_correct": 821,
  "parse_failed": 25,
  "outer_idx_out_of_range": 71,
  "target_not_assistant": 3
}

Format

Each JSONL row is compatible with ms-swift RLHF/DPO training and includes:

  • —id
  • —messages
  • —rejected_messages
  • —images
  • —audios
  • —videos
  • —rejected_images
  • —rejected_audios
  • —rejected_videos
  • —answer
  • —rollout_predicted_answer
  • —rollout_status
  • —first_error_index
  • —error_reason

Caveats

  • —This is derived preference data, not the full raw rollout/audit trace release.
  • —One rollout shard was excluded because it repeatedly timed out.
  • —The corrected branch contains Gemini-generated corrected sub-steps; tool observations were not re-executed and inserted into the chosen branch.
  • —OmniGAIA train/test share a media pool by benchmark design. The upstream SFT checkpoint verifies no test-set question text appears verbatim in training samples.

Intended use

Use this dataset to reproduce or inspect the DPO stage of ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO.

It is not intended as a general-purpose RLHF dataset outside OmniGAIA-style omni-modal tool-use agents.