CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tintin1027 /atomic-metrics-six-task-preferences Six-task benchmark inputs Seed 17. No demographic conditioning. Each task has shared train100.jsonl and test500.jsonl for Atomic Metrics, five judge variants, and learned baselines. Pair plans cover all 100 training rows once. Atomic Metrics extraction and BT/LR fitting use train100. Judges use the same test500. RM and WIMHF in the matched-data comparison use train100; rm_train_full is an explicitly separate expanded-data setting and must not be described as train100.… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-six-task-preferences.texttext-classification10K<n<100K0 likes222 downloads6d agoHugging Face02rlundqvist /ifeval-obf-rl-preferences IFEval Obfuscation — Full Preference Pairs (2023 constitution) Preference pairs over responses from a Wood-Labs eval-aware 49B organism (nemotron-nas / DeciLM), judged under the 2023 Claude constitution, for training reward models / DPO on verbalized evaluation-awareness (VEA). These are the FULL files the RMs actually trained on — not the earlier filtered subset. Files (DPO-ready) prefs_2023_leak_full.jsonl — 14,074 pairs. Judge saw the CoT + answer ("leak"… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/ifeval-obf-rl-preferences.texttext-generation10K<n<100K0 likes107 downloads28d agoHugging Face03spandyie /amadablam-dpo-preferences Ama Dablam DPO Preference Data Preference pairs used to DPO-tune Ama Dablam, a 322M trilingual (Nepali/Maithili/Bhojpuri) language model, across all three languages and three writing systems (Devanagari, IAST, phonetic romanization). See the technical report §9 for full methodology. Splits split rows purpose train 14,152 DPO Stage 2 preference-optimization training validation 744 preference-accuracy / forgetting evaluation warmup 3,203 Stage 1… See the full description on the dataset page: https://huggingface.co/datasets/spandyie/amadablam-dpo-preferences.texttext-generation10K<n<100K0 likes54 downloads20d agoHugging Face04ServiceNow-AI /Curriculum_DPO_preferences Curriculum DPO Preference Pairs This repository provides the curriculum DPO preference pairs used in the paper Curri-DPO, which explores enhancing model alignment through curriculum learning and ranked preferences. Datasets Ultrafeedback The Ultrafeedback dataset contains 64K preference pairs. We randomly sample 5K pairs and rank responses for each prompt, organizing them into three difficulty levels: easy, medium, and hard, based on response scores.… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/Curriculum_DPO_preferences.texttext-generation1K<n<10K6 likes37 downloads2y agoHugging Face05Hub-Ai /ptbr-human-preferences 🇧🇷 HUBX Human Preference Dataset (PT-BR) The largest Portuguese-Brazilian human preference dataset for RLHF/DPO training. 📊 Dataset Statistics Metric Value Total Annotations 314,757 Unique Tasks 450 Human Annotators ~600 Avg. Votes per Task ~699 Language Portuguese (Brazil) Domain Communication Quality & Tone 🎯 Why This Dataset? 🇧🇷 Native PT-BR: Collected from Brazilian Portuguese speakers - not translated 👥 Real Humans:… See the full description on the dataset page: https://huggingface.co/datasets/Hub-Ai/ptbr-human-preferences.texttext-classificationn<1K0 likes15 downloads10mo agoHugging Face06DebateLabKIT /argunauts-hirpo-preferences Argunauts HIRPO Preferences Preference pairs generated while training Argunaut models with HIRPO Online DPO. texttext-generation100K<n<1M0 likes14 downloads10mo agoHugging Face07lianghsun /ultrafeedback-binarized-preferences-cleaned-multilingualgated Dataset Card for ultrafeedback-binarized-preferences-cleaned-multilingual 本資料集是 argilla/ultrafeedback-binarized-preferences-cleaned 的多語言(含繁體中文 zh-tw)版本,每筆樣本保留原始來源、語言、對話與 chosen / rejected 回應,可作為繁中模型的 DPO 對齊資料。 Dataset Details Dataset Description 原始 UltraFeedback 是一個英文偏好資料集,本資料集將其翻譯/增廣為多語版本,並對譯文做品質清理,以利非英文(特別是繁中)DPO 訓練。 每筆樣本包含: source:原始來源(如 evol_instruct、flan_v2_p3 等)。 lang:語言代碼(如 zh-tw、en)。 conversations:human/gpt 對話結構,內容已翻譯為對應語言。 (以及對應的 chosen / rejected… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/ultrafeedback-binarized-preferences-cleaned-multilingual.texttext-generation1K<n<10K0 likes6 downloads5mo agoHugging Face08Amin-AQ /smollm2-dpo-preferencesgated DPO Preferences Dataset (Restricted Access) Access Policy (Restricted) This dataset repo is public with manual gated access. Only approved users (from lums.edu.pk) will be granted access. Intended Use Preference optimization / DPO experiments for model alignment. Research and controlled evaluation. Out-of-Scope Use Any harmful, abusive, or policy-violating application. Safety-critical deployment without additional safeguards. Files… See the full description on the dataset page: https://huggingface.co/datasets/Amin-AQ/smollm2-dpo-preferences.texttext-generation1K<n<10K0 likes4 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.