datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
概要
5種類のオープンモデルとQwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使って作成した、190854件の日本語合成Preferenceデータセットです。
以下、データセットの詳細です。
instructionには、Aratako/Magpie-Tanuki-8B-annotated-96kのinput_qualityがexcellentのものを利用
回答生成には、以下の5つのApache 2.0ライセンスのモデルを利用
weblab-GENIAC/Tanuki-8B-dpo-v1.0
team-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0-GPTQ-8bit
cyberagent/calm3-22b-chat
llm-jp/llm-jp-3-13b-instruct
Qwen/Qwen2.5-32B-Instruct-GPTQ-Int8… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k.dataset-tldr-preference-dpo
Dataset Card for dataset-tldr-preference-dpo
This dataset has been created with distilabel.
Dataset Summary
This is a dataset intended for training models using DPO/ORPO for the task of producing concise tl;dr summaries of machine learning datasets based on their dataset cards.
The dataset was created with distilabel. Each row of the dataset contains a dataset card which has been parsed to remove empty sections and placeholder text.
The instruction request… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/dataset-tldr-preference-dpo.MNLP_M1_Preference_dpo_dataset
M1 Preference Data for DPO
Dataset Description
This dataset contains processed M1 preference data for DPO training.
Created by: CS-552 Stochastic Parrots Team
Date: May 24, 2025
Version: 1.0
Number of examples: 17615
Dataset Source
This dataset is derived from the M1 preference data collected through interactions with large language models (like ChatGPT) for CS-552 (Modern Natural Language Processing) at EPFL. The preference data consists of… See the full description on the dataset page: https://huggingface.co/datasets/stochastic-parrots/MNLP_M1_Preference_dpo_dataset.judgment-consistency-preference-data
Dataset Card for judgment Consistency Preference Data
Dataset Description
This is a preference dataset designed to enhance the consistency of judgment in models when faced with disturbance, suitable for the DPO algorithm. It contains 2607 prompts sampled from arithmetic, commonsense, symbolic, and knowledge reasoning datasets, each accompanied by a pair of responses: one "chosen" response and one "rejected" response.
We design a dialogue scenario with one round of… See the full description on the dataset page: https://huggingface.co/datasets/NUSTM/judgment-consistency-preference-data.Self_Alignment_Preference-Dataset
Mistral Self-Alignment Preference Dataset
Warning: This dataset contains harmful and offensive data! Proceed with caution.
The Mistral Self-Alignment Preference Dataset was generated by Mistral 7b using the Anthropics Red Teaming Prompts dataset available at Hugging Face - Anthropics Red Teaming Prompts Dataset. The data generation process utilized the Preference Data Generation Notebook, which can be found here.
The purpose of this dataset is to facilitate self-alignment, as… See the full description on the dataset page: https://huggingface.co/datasets/August4293/Self_Alignment_Preference-Dataset.tulu-3-preference-data-with-distraction
Tulu-3 Preference Data with Distraction (Preference Data)
This dataset provides preference pairs (for DPO, IPO, ORPO, KTO, etc.) where prompts intentionally include distractor content (e.g., hidden instructions, puzzles, or extra tasks) to test and train models to ignore the distractor and solve the primary query. It is the preference companion to the SFT-only dataset groupfairnessllm/tulu-3-sft-with-distraction. The original data is derived from Tulu 3 dataset which contains coding… See the full description on the dataset page: https://huggingface.co/datasets/groupfairnessllm/tulu-3-preference-data-with-distraction.synthetic-preference-data
Synthetic Preference Data
A small, synthetically generated preference dataset intended for testing
RLHF / DPO training pipelines. Each example contains a prompt and two
responses — one correct (chosen), one subtly flawed (rejected).
Generation Procedure
Generator model: gpt-4o-mini (OpenAI)
Generation method: OpenAI Structured Outputs (response_format=PreferenceExample) — guarantees schema-valid JSON.
Prompt templates: 4 (factual, step-by-step reasoning, technical how-to… See the full description on the dataset page: https://huggingface.co/datasets/antony-bryan-3D2Y/synthetic-preference-data.
