datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
preference_dataset_mixture2_and_safe_pku
Copy from https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question… See the full description on the dataset page: https://huggingface.co/datasets/OpenRLHF/preference_dataset_mixture2_and_safe_pku.pair-preference-dataset-700K_standardpreference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question with this reward model and also any question about reward modeling, feel free to drop me an… See the full description on the dataset page: https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku.pair-preference-dataset-700K_subset-15-out-of-16_standardproactivity_preference_dataset
ProVoice study 1 — driver state, vehicle context and preferred Level of Autonomy
Driving-simulator data from the population data collection of the ProVoice /
ProActivity project (CARLA 0.10): 12 drivers × 2 sessions, ~20 Hz
multimodal driver-state and vehicle frames, and 1,446 driver-assigned
Level-of-Autonomy (LoA) labels stating how autonomously an in-vehicle
assistant should act on a given task. Drivers were prompted every 20 s about
two randomly drawn in-vehicle tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ProVoice-proactivity/proactivity_preference_dataset.math-preference-dataset
Dataset Card for math-preference-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/math-preference-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/math-preference-dataset.1004_ti2t_preference_dataset_supple_30kpreference_dataset_mix2
Dataset Card for "preference_dataset_mix2"
More Information needed
example-generate-preference-dataset
Dataset Card for example-preference-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/example-preference-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/example-generate-preference-dataset.pair-preference-dataset-700K_subset-9-out-of-10_standardpair-preference-dataset-700K_subset-7-out-of-8_standarddataset-for-annotation-v2-annotated
合成された質問と2つの応答文のペアに対して、日本語を母語とするチームメンバーが、好ましい応答文を人手でアノテーションしました
HelpSteer2-preferenceに習い、選好だけでなく選好の強度も[-3, 3]の範囲で付与しました
アノテーション時のメモ
ジャンルは以下
簡単な一般知識(wikipediaを読まずに回答できる系)
難しめの一般知識(wikipediaを読んだら回答できる系)
歴史上の出来事の論述
医療知識(応急処置系)
機械学習の課題と解決方法
化学式の解説
架空の物語生成
ロールプレイ
詩の創作
素因数分解や偶数奇数判定などの簡単な数学タスク
コーディングタスク
アルゴリズムやシステムのメリデメの解説
美術や思想についての論述
日本語の文法の解説
その他
LLMの定型文として登場するフレーズは、
「もちろんです」
「~も見逃せません」
「~も見過ごせません」
「総じて、」
「まず始めに、~さらに、~次に、~まとめると、」
データセットを目視で読み込んだ印象… See the full description on the dataset page: https://huggingface.co/datasets/preference-team/dataset-for-annotation-v2-annotated.pair-preference-dataset-700K_subset-3-of-4_standardpreference_dataset_mixture2_and_safe_pku150k
Dataset Card for "preference_dataset_mixture2_and_safe_pku150k"
More Information needed
preference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question with this reward model and also any question about reward modeling, feel free to drop me an… See the full description on the dataset page: https://huggingface.co/datasets/living-box/preference_dataset_mixture2_and_safe_pku.example-preference-dataset
Dataset Card for example-preference-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Yuriy81/example-preference-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Yuriy81/example-preference-dataset.0930_ta2t_preference_dataset_20Kpair-preference-dataset-700K_subset-2-of-2_gemma-2b-it_1-of-2MNLP_M1_Preference_dpo_dataset
M1 Preference Data for DPO
Dataset Description
This dataset contains processed M1 preference data for DPO training.
Created by: CS-552 Stochastic Parrots Team
Date: May 24, 2025
Version: 1.0
Number of examples: 17615
Dataset Source
This dataset is derived from the M1 preference data collected through interactions with large language models (like ChatGPT) for CS-552 (Modern Natural Language Processing) at EPFL. The preference data consists of… See the full description on the dataset page: https://huggingface.co/datasets/stochastic-parrots/MNLP_M1_Preference_dpo_dataset.pair-preference-dataset-700K_subset-2-of-3_standardpair-preference-dataset-700K_subset-4-of-4_gemma-2b_1of4_iter1_conf-0.8_bs128_lr1e-5_conf-0.8pair-preference-dataset-700K_subset-2-of-4_gemma-2b_1of4_iter3_conf-0.8_bs128_lr1e-5_conf-0.8pair-preference-dataset-700K_subset-2-of-3_soup_conf-0.9pair-preference-dataset-700K_subset-4-of-4_standardpair-preference-dataset-700K_subset-1-of-4_standardpair-preference-dataset-700K_subset-1-of-3_standardpair-preference-dataset-700K_subset-2-of-3_soup_conf-0.6pair-preference-dataset-700K_subset-2-of-2_standardpair-preference-dataset-700K_subset-2-of-3_5e-6_conf-0.9R3-Dataset-15K-Preference-Only-v1.1
