datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset
Rapidata Image Generation Preference Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.preference-datasets-tulupreference_dataset_mixture2_and_safe_pku
Copy from https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question… See the full description on the dataset page: https://huggingface.co/datasets/OpenRLHF/preference_dataset_mixture2_and_safe_pku.pair_preference_model_dataset_add_prefix_to_win_rate0.1_rrm_0p2pair_preference_model_dataset_add_emoji_to_win_rate0.1_rrm_newpair-preference-dataset-700K_standardpreference_dataset_mixture2_and_safe_pku
Reward Model Overview
This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...): https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
Model Details
If you have any question with this reward model and also any question about reward modeling, feel free to drop me an… See the full description on the dataset page: https://huggingface.co/datasets/weqweasdas/preference_dataset_mixture2_and_safe_pku.universal-preference-hijacking-datasets
Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
Figure 1: Examples of Phi, which can hijack MLLM's preference toward the image.
Figure 2: Example of a universal hijacking perturbation, which can be transferred across different images.
This dataset is used to train and evaluate the universal hijacking perturbations in the paper "Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time", accepted at EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/yflantmy/universal-preference-hijacking-datasets.Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k
概要
5種類のオープンモデルとQwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使って作成した、190854件の日本語合成Preferenceデータセットです。
以下、データセットの詳細です。
instructionには、Aratako/Magpie-Tanuki-8B-annotated-96kのinput_qualityがexcellentのものを利用
回答生成には、以下の5つのApache 2.0ライセンスのモデルを利用
weblab-GENIAC/Tanuki-8B-dpo-v1.0
team-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0-GPTQ-8bit
cyberagent/calm3-22b-chat
llm-jp/llm-jp-3-13b-instruct
Qwen/Qwen2.5-32B-Instruct-GPTQ-Int8… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k.pair_preference_model_dataset_gemma2_2b_rrm_0p2pair-preference-dataset-700K_subset-15-out-of-16_standardproactivity_preference_dataset
ProVoice study 1 — driver state, vehicle context and preferred Level of Autonomy
Driving-simulator data from the population data collection of the ProVoice /
ProActivity project (CARLA 0.10): 12 drivers × 2 sessions, ~20 Hz
multimodal driver-state and vehicle frames, and 1,446 driver-assigned
Level-of-Autonomy (LoA) labels stating how autonomously an in-vehicle
assistant should act on a given task. Drivers were prompted every 20 s about
two randomly drawn in-vehicle tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ProVoice-proactivity/proactivity_preference_dataset.pair_preference_model_dataset_add_prefix_to_win_rate0.1_rrm_newmath-preference-dataset
Dataset Card for math-preference-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/math-preference-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/math-preference-dataset.preference_dataset-standard_format-v2.21004_ti2t_preference_dataset_supple_30kPreference_Dataset_Merged
Dataset Overview
This Dataset consists of the following open-sourced preference dataset
Arena Human Preference
Anthropic HH
MT-Bench Human Judgement
Ultra Feedback
Tulu3 Preference Dataset
Skywork-Reward-Preference-80K-v0.2
Cleaning
Cleaning Method 1: Only keep the following Language using FastText language detection(EN/DE/ES/ZH/IT/JA/FR)
Cleaning Method 2: Remove duplicates to ensure each prompt appears only once
Cleaning Method 3: Remove datasets where… See the full description on the dataset page: https://huggingface.co/datasets/yufan/Preference_Dataset_Merged.pair_preference_model_dataset_add_prefix_to_win_rate0.1_rrmpreference_dataset_mix2
Dataset Card for "preference_dataset_mix2"
More Information needed
example-generate-preference-dataset
Dataset Card for example-preference-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/example-preference-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/example-generate-preference-dataset.dataset-tldr-preference
Dataset Card for dataset-tldr-preference
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/davanstrien/dataset-tldr-preference/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/dataset-tldr-preference.combined-preference-datasetCombined preference dataset. All examples are binarized and standardized for tokenizer.apply_chat_template().
Source datasets:
openbmb/UltraFeedback
coseal/CodeUltraFeedback
nvidia/HelpSteer2
PKU-Alignment/PKU-SafeRLHF
argilla/Capybara-Preferences-Filtered
argilla/distilabel-intel-orca-dpo-pairs
argilla/distilabel-math-preference-dpo
stanfordnlp/SHP
1002_preference_dataset_30k_promptpharma-preference-dataset
Pharma DPO Preference Dataset
Pharmaceutical domain preference dataset used for
Direct Preference Optimization (DPO) — Stage 3 of the pharma TinyLlama
fine-tuning pipeline.
Format
Each JSONL record contains 3 fields:
{
"prompt": "### Instruction:\nExplain the mechanism of metformin.\n\n### Response:\n",
"chosen": "Metformin primarily works by ...",
"rejected": "Metformin is a drug that ..."
}
prompt — Alpaca-style instruction prompt (same format as… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-preference-dataset.pair_preference_model_datasetpair-preference-dataset-700K_subset-9-out-of-10_standardpair_preference_model_dataset_add_prefix_to_win_rate0.1_rrm_post_inference_4p5mto5mexample-preference-dataset2
Dataset Card for example-preference-dataset2
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/ashercn97/example-preference-dataset2/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/ashercn97/example-preference-dataset2.pair_preference_model_dataset_add_prefix_to_win_rate0.1tldr_preference_dataset
Dataset Card for "tldr_preference_dataset"
More Information needed
