CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01argilla /ultrafeedback-binarized-preferences-cleaned UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md. Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.tabulartext-generation10K<n<100K165 likes27k downloads3y agoHugging Face02argilla /ultrafeedback-binarized-preferences-cleaned-kto UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) KTO A KTO signal transformed version of the highly loved UltraFeedback Binarized Preferences Cleaned, the preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned-kto.texttext-generation100K<n<1M10 likes16k downloads3y agoHugging Face03euclaise /WritingPrompts_preferences Dataset Card for "WritingPrompts_preferences" Human preference data from r/WritingPrompts texttext-generation100K<n<1M13 likes1.2k downloads3y agoHugging Face04argilla /Capybara-Preferences Dataset Card for Capybara-Preferences This dataset has been created with distilabel. Dataset Summary This dataset is built on top of LDJnr/Capybara, in order to generate a preference dataset out of an instruction-following dataset. This is done by keeping the conversations in the column conversation but splitting the last assistant turn from it, so that the conversation contains all the turns up until the last user's turn, so that it can be reused… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Capybara-Preferences.tabulartext-generation10K<n<100K47 likes436 downloads2y agoHugging Face05tintin1027 /atomic-metrics-six-task-preferences Six-task benchmark inputs Seed 17. No demographic conditioning. Each task has shared train100.jsonl and test500.jsonl for Atomic Metrics, five judge variants, and learned baselines. Pair plans cover all 100 training rows once. Atomic Metrics extraction and BT/LR fitting use train100. Judges use the same test500. RM and WIMHF in the matched-data comparison use train100; rm_train_full is an explicitly separate expanded-data setting and must not be described as train100.… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-six-task-preferences.texttext-classification10K<n<100K0 likes222 downloads7d agoHugging Face06surrey-nlp /dialect-preferences DiaLLM — Pooled Preference Dataset (Implicit Thread) Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 45,690 preference pairs, pooling all three variety-specific sets (Australian, Northern British, Indian) without variety targeting. Used for implicit-thread DPO training, where the three varieties are pooled rather than targeted individually, preserving the variety-agnostic objective of that thread.… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/dialect-preferences.tabulartext-generation10K<n<100K0 likes191 downloads1mo agoHugging Face07argilla /ultrafeedback-multi-binarized-preferences-cleaned UltraFeedback - Multi-Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences-cleaned, and has been created to explore whether DPO fine-tuning with more than one rejection per chosen response helps the model perform better in the AlpacaEval, MT-Bench, and LM Eval Harness benchmarks. Read more about Argilla's approach towards UltraFeedback binarization at… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-multi-binarized-preferences-cleaned.tabulartext-generation100K<n<1M7 likes164 downloads3y agoHugging Face08Polygl0t /gigaverbo-v2-preferences GigaVerbo-v2 Preferences: A Hybrid-Reasoning Portuguese Preference Dataset Dataset Summary GigaVerbo-v2 Preferences is a preference dataset designed for Direct Preference Optimization (DPO) and other direct alignment algorithms. The dataset comprises approximately 27.8 million tokens across 28,437 preference pairs, organized into 4 distinct subsets covering both quality-focused and safety-focused alignment. It is entirely composed of high-quality, LLM-generated data… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/gigaverbo-v2-preferences.tabulartext-generation10K<n<100K0 likes116 downloads7mo agoHugging Face09vicgalle /creative-rubrics-preferences creative-rubrics-preferences 🎏 A dataset of creative responses using GPT-4.5, o3-mini and DeepSeek-R1. This dataset contains several prompts seeking creative and diverse answers (like writing movie reviews, short stories, etc), and the style of the responses has been enhanced by prompting the model with custom rubrics that seek different creative styles. This dataset was used in the paper Configurable Preference Tuning with Rubric-Guided Synthetic Data. Code:… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/creative-rubrics-preferences.texttext-generationn<1K4 likes113 downloads1y agoHugging Face10rlundqvist /ifeval-obf-rl-preferences IFEval Obfuscation — Full Preference Pairs (2023 constitution) Preference pairs over responses from a Wood-Labs eval-aware 49B organism (nemotron-nas / DeciLM), judged under the 2023 Claude constitution, for training reward models / DPO on verbalized evaluation-awareness (VEA). These are the FULL files the RMs actually trained on — not the earlier filtered subset. Files (DPO-ready) prefs_2023_leak_full.jsonl — 14,074 pairs. Judge saw the CoT + answer ("leak"… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/ifeval-obf-rl-preferences.texttext-generation10K<n<100K0 likes107 downloads28d agoHugging Face11argilla /Capybara-Preferences-Filtered Dataset Card for Capybara-Preferences-Filtered This dataset has been created with distilabel, plus some extra post-processing steps described below. Dataset Summary This dataset is built on top of argilla/Capybara-Preferences, but applies a further in detail filtering. The filtering approach has been proposed and shared by @LDJnr, and applies the following: Remove responses from the assistant, not only in the last turn, but also in intermediate… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Capybara-Preferences-Filtered.tabulartext-generation10K<n<100K10 likes102 downloads2y agoHugging Face12kristaller486 /wikisource_preferences_ru Wikisource Preferences [Russian] Датасет для оптимизации предпочтений. chosen тексты брались из kristaller486/wikisource-creative-ru, а rejected генерировались разнообразными LLM по сгенерированным промптам. Шаблон для DPO: axolotl chat_template.default Модели для генерации rejected семплов: google/gemma-3-27b-it gpt-4.1-mini gpt-4.1-nano gpt-4.1 gemini-2.0-flash Qwen/Qwen3-14B-FP8 (without reasoning) Moraliane/SAINEMO-reMIX (fp6-llm quantization) deepseek-v3-0324 (api)… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/wikisource_preferences_ru.texttext-generation10K<n<100K0 likes79 downloads1y agoHugging Face13wassname /ethics_expression_preferencesReformat of the ethics dataset to be in DPO format, in this expression format Examples: subset commonsense before 13910 {'label': 1, 'input': "I went to the principal's office to change my records before going to a different school."} after 1513 {'prompt': 'I won my match by ', 'chosen': 'making a lot of good shots.', 'rejected': "moving my ball when the other player wasn't looking."} subset utilitarianism before 13737 {'baseline': 'I forgot to wear my virus-proof mask when I visited… See the full description on the dataset page: https://huggingface.co/datasets/wassname/ethics_expression_preferences.texttext-generation10K<n<100K3 likes74 downloads1y agoHugging Face14saitejaalasyam /grounded-qa-preferences Grounded QA preferences Preference pairs for a small RLHF stack. Each row is a passage, a question, a preferred answer, and a rejected answer. The questions, answer spans, and unanswerable labels come from SQuAD 2.0 (Rajpurkar et al.). This dataset does not add new human rankings. A fixed rule turns those annotations into Bradley-Terry pairs: pair_type When Chosen Rejected wrong_span The passage answers the question The gold span A different short span from the same… See the full description on the dataset page: https://huggingface.co/datasets/saitejaalasyam/grounded-qa-preferences.texttext-generation1K<n<10K1 likes72 downloads1d agoHugging Face15csarron /argilla-ultrafeedback-binarized-preferences-cleaned UltraFeedback (Cleaned) This dataset combines the train split of argilla/ultrafeedback-binarized-preferences-cleaned, and test split of HuggingFaceH4/ultrafeedback_binarized. texttext-generation10K<n<100K1 likes70 downloads2y agoHugging Face16spandyie /amadablam-dpo-preferences Ama Dablam DPO Preference Data Preference pairs used to DPO-tune Ama Dablam, a 322M trilingual (Nepali/Maithili/Bhojpuri) language model, across all three languages and three writing systems (Devanagari, IAST, phonetic romanization). See the technical report §9 for full methodology. Splits split rows purpose train 14,152 DPO Stage 2 preference-optimization training validation 744 preference-accuracy / forgetting evaluation warmup 3,203 Stage 1… See the full description on the dataset page: https://huggingface.co/datasets/spandyie/amadablam-dpo-preferences.texttext-generation10K<n<100K0 likes54 downloads20d agoHugging Face17pharaouk /ultrafeedback-binarized-preferences-cleaned UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md. Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/ultrafeedback-binarized-preferences-cleaned.tabulartext-generation10K<n<100K0 likes44 downloads2y agoHugging Face18anakin87 /gemma-vs-gemma-preferences 💎🆚💎 Gemma vs Gemma Preferences This dataset contains on-policy collected preferences generated using anakin87/gemma-2-2b-ita-sft. ⚠️ While this dataset may be valuable for didactic purposes, it is not recommended for training a model using Preference Tuning due to the following reasons: The training would be off-policy for your model. The dataset was generated with gemma-2-2b-ita-sft, a small model for Italian. Motivation While DPO (Direct Preference… See the full description on the dataset page: https://huggingface.co/datasets/anakin87/gemma-vs-gemma-preferences.texttext-generation10K<n<100K0 likes44 downloads2y agoHugging Face19wassname /truthful_qa_preferencestexttext-classificationn<1K0 likes43 downloads2y agoHugging Face20wassname /ethics_qna_preferencesReformat of the ethics dataset to be in DPO format, in this expression format Examples: subset virtue {'prompt': 'Sentence: Jill wore a new style of distinctive clothes each six months for decades.\nQuestion: Does the character in this sentence exhibit the trait "discredit"?\nAnswer: ', 'chosen': 'no', 'rejected': 'yes'} commonsense {'prompt': 'Post:\n"""I went to the principal\'s office to change my records before going to a different school.""""\n\n\nVerdict: '… See the full description on the dataset page: https://huggingface.co/datasets/wassname/ethics_qna_preferences.textquestion-answering100K<n<1M1 likes42 downloads1y agoHugging Face21kixlab /DiscoverLLM-multiturn-preferences DiscoverLLM: Multi-turn Preference Dataset Multi-turn dialogue data with scored candidate completions, produced by best-of-N synthesis over the DiscoverLLM user simulator (paper · project page). Each example is a single turn of a simulated user–assistant conversation with one of several candidate assistant responses and an associated reward score, intended for offline DPO / GRPO / reward-model training. Configs Config Rows Task creative_writing 3,052… See the full description on the dataset page: https://huggingface.co/datasets/kixlab/DiscoverLLM-multiturn-preferences.tabulartext-generation1K<n<10K3 likes39 downloads4mo agoHugging Face22ServiceNow-AI /Curriculum_DPO_preferences Curriculum DPO Preference Pairs This repository provides the curriculum DPO preference pairs used in the paper Curri-DPO, which explores enhancing model alignment through curriculum learning and ranked preferences. Datasets Ultrafeedback The Ultrafeedback dataset contains 64K preference pairs. We randomly sample 5K pairs and rank responses for each prompt, organizing them into three difficulty levels: easy, medium, and hard, based on response scores.… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/Curriculum_DPO_preferences.texttext-generation1K<n<10K6 likes37 downloads2y agoHugging Face23alvarobartt /openhermes-preferences-coding Dataset Card for OpenHermes Preferences - Coding This dataset is a subset from argilla/OpenHermesPreferences, only keeping the preferences of the source coding, and removing all the columns besides the chosen and rejected ones, that come in OpenAI chat formatting, so that's easier to fine-tune a model using tools like: huggingface/alignment-handbook or axolotl, among others. Reference argilla/OpenHermesPreferences dataset created as a collaborative effort between… See the full description on the dataset page: https://huggingface.co/datasets/alvarobartt/openhermes-preferences-coding.texttext-generation1K<n<10K5 likes32 downloads3y agoHugging Face24alvarobartt /openhermes-preferences-metamath Dataset Card for OpenHermes Preferences - MetaMath This dataset is a subset from argilla/OpenHermesPreferences, only keeping the preferences of metamath, and removing all the columns besides the chosen and rejected ones, that come in OpenAI chat formatting, so that's easier to fine-tune a model using tools like: huggingface/alignment-handbook or axolotl, among others. Reference argilla/OpenHermesPreferences dataset created as a collaborative effort between Argilla and… See the full description on the dataset page: https://huggingface.co/datasets/alvarobartt/openhermes-preferences-metamath.texttext-generation10K<n<100K4 likes28 downloads3y agoHugging Face25Heriot-WattUniversity /gbv-cs-binary-preferencestexttext-generation1K<n<10K0 likes28 downloads1y agoHugging Face26hallinh /Enterprise-RLHF-Preferences-10k-Sample 🏆 Enterprise RLHF Preference Dataset (10k Sample) ⚠️ RESEARCH & EVALUATION ONLY ⚠️ This is a 10,000-sample preview of the full 60k Enterprise Corpus. For the full commercial license and access to the complete dataset, please contact: [ alinmatei.dev@gmail.com] 📖 Overview This dataset represents a premium corpus for Reinforcement Learning from Human Feedback (RLHF) and Reward Model (RM) training. Unlike standard web-scraped datasets, this corpus focuses on… See the full description on the dataset page: https://huggingface.co/datasets/hallinh/Enterprise-RLHF-Preferences-10k-Sample.textreinforcement-learning10K<n<100K0 likes28 downloads9mo agoHugging Face27furquan /dialectic-preferences-bias-aae-sae-parallel Dialectic Preferences Bias Dataset Dataset Description Overview This dataset is part of a research study examining dialectic preference bias in Large Language Models (LLMs). It contains paired sentences in African American English (AAE) and Standard American English (SAE), used to analyze potential biases in language models' treatment of different dialects. The dataset contains two columns: african_american_english: Text samples in African American English… See the full description on the dataset page: https://huggingface.co/datasets/furquan/dialectic-preferences-bias-aae-sae-parallel.texttext-classification1K<n<10K0 likes23 downloads2y agoHugging Face28victor203 /fitness-preferences Fitness Preferences Dataset for RLHF Dataset Description This dataset contains human preference pairs for fitness and exercise-related questions, designed for training reward models and fine-tuning language models with RLHF (Reinforcement Learning from Human Feedback). Dataset Summary Total Size: 54 preference pairs Domain: Fitness, exercise, nutrition, and health Task: Preference learning for fitness advice generation Language: English License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/victor203/fitness-preferences.texttext-generationn<1K0 likes23 downloads8mo agoHugging Face29NordosoftOy /innoduel-rlhf-real-world-human-preferences-sample Real-World Human Pairwise Preferences — Public Sample 📦 This is a free, public sample of a commercial dataset. It contains 1,350 rows curated for inspection. The full dataset has 1.5 million human pairwise-preference decisions. Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi. Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.tabulartext-generation1K<n<10K0 likes23 downloads1mo agoHugging Face30Hub-Ai /ptbr-human-preferences 🇧🇷 HUBX Human Preference Dataset (PT-BR) The largest Portuguese-Brazilian human preference dataset for RLHF/DPO training. 📊 Dataset Statistics Metric Value Total Annotations 314,757 Unique Tasks 450 Human Annotators ~600 Avg. Votes per Task ~699 Language Portuguese (Brazil) Domain Communication Quality & Tone 🎯 Why This Dataset? 🇧🇷 Native PT-BR: Collected from Brazilian Portuguese speakers - not translated 👥 Real Humans:… See the full description on the dataset page: https://huggingface.co/datasets/Hub-Ai/ptbr-human-preferences.texttext-classificationn<1K0 likes15 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.