datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-instruct-gptj-pairwisemsmarco-llm-reranking-pairwisepairwise_preferencesv2oasst1_pairwise_rlhf_reward
Dataset Card for "oasst1_pairwise_rlhf_reward"
OASST1 dataset preprocessed for reward modeling:
import pandas as pd
from datasets import load_dataset,concatenate_datasets, Dataset, DatasetDict
import numpy as np
dataset = load_dataset("OpenAssistant/oasst1")
df=concatenate_datasets(list(dataset.values())).to_pandas()
m2t=df.set_index("message_id")['text'].to_dict()
m2r=df.set_index("message_id")['role'].to_dict()
m2p=df.set_index('message_id')['parent_id'].to_dict()… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/oasst1_pairwise_rlhf_reward.red_teaming_reward_modeling_pairwise
Dataset Card for "red_teaming_reward_modeling_pairwise"
More Information needed
fusion-pairwise-evals-test-time-scaling
Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N
Content
This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares CommandA against gemini-2.5-pro in 2 settings:
Test-time scaling with Fusion : 5 samples are generated from CommandA, then fused with CommandA into one completion and compared to a single completion from gemini-2.5-pro
Test-time scaling with BoN : 5 samples are generated from… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-test-time-scaling.fusion-pairwise-evals-finetuned
Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N
Content
This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares 2 models against gemini-2.5-flash:
Fusion: is the 111B model finetuned on synthetic data generated with Fusion from 5 teachers
BoN: is the 111B model finetuned on synthetic data generated with BoN from 5 teachers
Each model’s outputs are compared in pairs with the respective… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-finetuned.ToolPref-Pairwise-30K
ToolPref-Pairwise-30K
[Paper] |
[Model] |
[Benchmark] |
[Code]
💡 Summary
This dataset is a part of ToolRM: Towards Agentic Tool-Use Reward Modeling. It comprises 30,000 preference annotations in agentic tool-use scenarios and was used to train the ToolRM model series.
🌟 Overview
ToolRM is a family of lightweight generative and discriminative reward models tailored for agentic tool-use scenarios. To build these models, we propose a novel pipeline… See the full description on the dataset page: https://huggingface.co/datasets/RioLee/ToolPref-Pairwise-30K.davinci-pairwise-tokenized
Dataset Card for "davinci-pairwise-tokenized"
More Information needed
sidewalk_ballet_pairwisered_teaming_reward_modeling_pairwise_no_as_an_ai
Dataset Card for "red_teaming_reward_modeling_pairwise_no_as_an_ai"
More Information needed
deja-vu-pairwise-evals
Automatic pairwise preference evaluations for "Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation"
Content
This data contains pairwise automatic win-rate evaluations for 2 benchmarks.
Outputs and judge decisions for the m-ArenaHard benchmark for sampled generations (5 each) from Aya Expanse 8B and Qwen2.5 7B Instruct.
Original and roundtrip-translated prompts (by NLLB 3.3B, Aya Expanse 32B, Google Translate, Command A), outputs and… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/deja-vu-pairwise-evals.qwen3-omni-pairwise-video-infer
Qwen3-Omni Pairwise Video Inference / Evaluation
Pairwise audio-video preference evaluation data for Qwen3-Omni models.
Each sample compares two generated videos (with audio) against a text caption and human/Gemini labels.
Source path on cluster: /inspire/hdd/project/autoregressive-video-generation/public/hym/data/final_infer
Upload snapshot: 2026-06-12 10:45 UTC
Repository layout
Contents of final_infer are uploaded to the dataset repo root:
.cache/
easy500/ —… See the full description on the dataset page: https://huggingface.co/datasets/YinmingHuang/qwen3-omni-pairwise-video-infer.sharegpt_reward_modeling_pairwise_no_as_an_ai
Dataset Card for "sharegpt_reward_modeling_pairwise_no_as_an_ai"
More Information needed
gsm8k_train_pairwise
Dataset Card for "gsm8k_train_pairwise"
More Information needed
pairwise-poisson-algebras
Pairwise Poisson Algebras: Neural Networks vs Physics
Dataset Description
This dataset contains the first systematic computation of pairwise Poisson bracket Lie algebras for both neural network training dynamics and physical N-body systems. SGD with momentum is a Hamiltonian system; the pairwise interactions between weight layers generate a Lie algebra — and we discover that neural networks produce richer algebraic structures than any physical system.
Neural… See the full description on the dataset page: https://huggingface.co/datasets/bshepp/pairwise-poisson-algebras.davinci-pairwise-medium
Dataset Card for "davinci-pairwise-medium"
More Information needed
orm-pairwise-preference-pairs
Pairwise Outcome Reward Model (ORM)
A Robust Preference Learning Model for Agentic Reasoning Systems
📋 Model Description
This is a Pairwise Outcome Reward Model (ORM) designed for agentic reasoning systems. The model learns to rank reasoning traces through relative preference judgments rather than absolute quality scores, achieving superior stability and reproducibility compared to traditional pointwise approaches.
Key Achievements:
✅ 96.3% pairwise accuracy with… See the full description on the dataset page: https://huggingface.co/datasets/LossFunctionLover/orm-pairwise-preference-pairs.gpteacher_reward_modeling_pairwise
Dataset Card for "gpteacher_reward_modeling_pairwise"
More Information needed
mirage-bench-pairwise-judgments
MIRAGE-Bench Pairwise Judgments
Win matrix per language from nthakur/mirage-bench-pairwise-judgments. Each cell (row, col) shows the win rate of the row model against the col model, computed as wins / total_comparisons × 100%. Ties are counted as 0.5 wins for each side. Each row in the dataset is treated as an independent outcome.
Arabic (ar)
Win Matrix — cell (row, col) = wins of row model vs col model out of 100 pairwise comparisons (ties = 0.5).
Diagonal is -.… See the full description on the dataset page: https://huggingface.co/datasets/nthakur/mirage-bench-pairwise-judgments.davinci-pairwise-filtered
Dataset Card for "davinci-pairwise-filtered"
More Information needed
SHP_pairwiselicense-pairwise-hf
Deprecated — moved to midah/hf-dataset-licenses
This repository is deprecated as of 2026-05-23.
All pairwise comparison data previously hosted here has been migrated to the
canonical license analysis dataset:
https://huggingface.co/datasets/midah/hf-dataset-licenses
The canonical repo contains everything that was here, plus:
The license corpus (corpus config, 747 licenses with full text + metadata)
Feature extractions (features_v3_* configs, schema v3)
SPDX-747 pairwise data… See the full description on the dataset page: https://huggingface.co/datasets/midah/license-pairwise-hf.davinci-vs-lit-pairwise
Dataset Card for "davinci-vs-lit-pairwise"
More Information needed
llm-metric-ace-code-pairwisesharegpt_reward_modeling_pairwise
Dataset Card for "sharegpt_reward_modeling_pairwise"
More Information needed
Design2Code_human_eval_pairwiseFind more details in our paper.
hh-rlhf-pairwise
Dataset Card for "hh-rlhf-pairwise"
More Information needed
creative-writing-pairwise-critic-freeformllm-metric-ace-code-pairwise-new
