datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
atomic-metrics-six-task-preferences
Six-task benchmark inputs
Seed 17. No demographic conditioning. Each task has shared train100.jsonl and test500.jsonl for Atomic Metrics, five judge variants, and learned baselines. Pair plans cover all 100 training rows once. Atomic Metrics extraction and BT/LR fitting use train100. Judges use the same test500. RM and WIMHF in the matched-data comparison use train100; rm_train_full is an explicitly separate expanded-data setting and must not be described as train100.… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-six-task-preferences.ifeval-obf-rl-preferences
IFEval Obfuscation — Full Preference Pairs (2023 constitution)
Preference pairs over responses from a Wood-Labs eval-aware 49B organism (nemotron-nas / DeciLM),
judged under the 2023 Claude constitution, for training reward models / DPO on verbalized
evaluation-awareness (VEA). These are the FULL files the RMs actually trained on — not the
earlier filtered subset.
Files (DPO-ready)
prefs_2023_leak_full.jsonl — 14,074 pairs. Judge saw the CoT + answer ("leak"… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/ifeval-obf-rl-preferences.amadablam-dpo-preferences
Ama Dablam DPO Preference Data
Preference pairs used to DPO-tune Ama Dablam,
a 322M trilingual (Nepali/Maithili/Bhojpuri) language model, across all three languages
and three writing systems (Devanagari, IAST, phonetic romanization). See the
technical report §9 for full
methodology.
Splits
split
rows
purpose
train
14,152
DPO Stage 2 preference-optimization training
validation
744
preference-accuracy / forgetting evaluation
warmup
3,203
Stage 1… See the full description on the dataset page: https://huggingface.co/datasets/spandyie/amadablam-dpo-preferences.Curriculum_DPO_preferences
Curriculum DPO Preference Pairs
This repository provides the curriculum DPO preference pairs used in the paper Curri-DPO, which explores enhancing model alignment through curriculum learning and ranked preferences.
Datasets
Ultrafeedback
The Ultrafeedback dataset contains 64K preference pairs. We randomly sample 5K pairs and rank responses for each prompt, organizing them into three difficulty levels: easy, medium, and hard, based on response scores.… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/Curriculum_DPO_preferences.ptbr-human-preferences
🇧🇷 HUBX Human Preference Dataset (PT-BR)
The largest Portuguese-Brazilian human preference dataset for RLHF/DPO training.
📊 Dataset Statistics
Metric
Value
Total Annotations
314,757
Unique Tasks
450
Human Annotators
~600
Avg. Votes per Task
~699
Language
Portuguese (Brazil)
Domain
Communication Quality & Tone
🎯 Why This Dataset?
🇧🇷 Native PT-BR: Collected from Brazilian Portuguese speakers - not translated
👥 Real Humans:… See the full description on the dataset page: https://huggingface.co/datasets/Hub-Ai/ptbr-human-preferences.argunauts-hirpo-preferences
Argunauts HIRPO Preferences
Preference pairs generated while training Argunaut models with HIRPO Online DPO.
ultrafeedback-binarized-preferences-cleaned-multilingual
Dataset Card for ultrafeedback-binarized-preferences-cleaned-multilingual
本資料集是 argilla/ultrafeedback-binarized-preferences-cleaned 的多語言(含繁體中文 zh-tw)版本,每筆樣本保留原始來源、語言、對話與 chosen / rejected 回應,可作為繁中模型的 DPO 對齊資料。
Dataset Details
Dataset Description
原始 UltraFeedback 是一個英文偏好資料集,本資料集將其翻譯/增廣為多語版本,並對譯文做品質清理,以利非英文(特別是繁中)DPO 訓練。
每筆樣本包含:
source:原始來源(如 evol_instruct、flan_v2_p3 等)。
lang:語言代碼(如 zh-tw、en)。
conversations:human/gpt 對話結構,內容已翻譯為對應語言。
(以及對應的 chosen / rejected… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/ultrafeedback-binarized-preferences-cleaned-multilingual.smollm2-dpo-preferences
DPO Preferences Dataset (Restricted Access)
Access Policy (Restricted)
This dataset repo is public with manual gated access.
Only approved users (from lums.edu.pk) will be granted access.
Intended Use
Preference optimization / DPO experiments for model alignment.
Research and controlled evaluation.
Out-of-Scope Use
Any harmful, abusive, or policy-violating application.
Safety-critical deployment without additional safeguards.
Files… See the full description on the dataset page: https://huggingface.co/datasets/Amin-AQ/smollm2-dpo-preferences.
