datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fusion-pairwise-evals-test-time-scaling
Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N
Content
This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares CommandA against gemini-2.5-pro in 2 settings:
Test-time scaling with Fusion : 5 samples are generated from CommandA, then fused with CommandA into one completion and compared to a single completion from gemini-2.5-pro
Test-time scaling with BoN : 5 samples are generated from… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-test-time-scaling.fusion-pairwise-evals-finetuned
Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N
Content
This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares 2 models against gemini-2.5-flash:
Fusion: is the 111B model finetuned on synthetic data generated with Fusion from 5 teachers
BoN: is the 111B model finetuned on synthetic data generated with BoN from 5 teachers
Each model’s outputs are compared in pairs with the respective… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-finetuned.orm-pairwise-preference-pairs
Pairwise Outcome Reward Model (ORM)
A Robust Preference Learning Model for Agentic Reasoning Systems
📋 Model Description
This is a Pairwise Outcome Reward Model (ORM) designed for agentic reasoning systems. The model learns to rank reasoning traces through relative preference judgments rather than absolute quality scores, achieving superior stability and reproducibility compared to traditional pointwise approaches.
Key Achievements:
✅ 96.3% pairwise accuracy with… See the full description on the dataset page: https://huggingface.co/datasets/LossFunctionLover/orm-pairwise-preference-pairs.pairwise_compare_all_dpo_llama_3_70b_datasynthetic-instruct-gptj-pairwise-ja
Dahoas/synthetic-instruct-gptj-pairwise-ja
Dahoas/synthetic-instruct-gptj-pairwiseの和訳
tasksource_oasst2_pairwise_rlhf_reward-PreferenceShareGPTpairwise_compare_top1_and_SPINS_llama_3_70b_dpo_datapairwise_only_compare_rank_1_to_all_llama_3_70b_dpo_dataqwen3-omni-pairwise-video-train
Qwen3-Omni Pairwise Video Inference / Evaluation
Pairwise audio-video preference evaluation data for Qwen3-Omni models.
Each sample compares two generated videos (with audio) against a text caption and human/Gemini labels.
Source path on cluster: /inspire/hdd/project/autoregressive-video-generation/public/hym/data/final_train
Upload snapshot: 2026-06-12 10:46 UTC
Repository layout
Contents of final_infer are uploaded to the dataset repo root:
.cache/
ovi_davinci/… See the full description on the dataset page: https://huggingface.co/datasets/YinmingHuang/qwen3-omni-pairwise-video-train.polititune-tankie-pairwise
