datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RLHF-V-Dataset
Dataset Card for RLHF-V-Dataset
Project Page | Paper | GitHub
Updates
[2024.05.28] 📃 Our RLAIF-V paper is accesible at arxiv now!
[2024.05.20] 🎉 We release a new feedback dataset, RLAIF-V-Dataset, which is a large-scale diverse-task multimodal feedback dataset constructed using open-source models. You can download the corresponding dataset and models (7B, 12B) now!
[2024.04.11] 🔥 Our data is used in MiniCPM-V 2.0, an end-side multimodal large language model that… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLHF-V-Dataset.RLHF-VBorrowed from: https://huggingface.co/datasets/openbmb/RLHF-V-Dataset
You can use it in LLaMA Factory by specifying dataset: rlhf_v.
generative-rlhf-training-data
Generative RLHF-V Online Interaction Training Data
This repository contains the exact 4,053-example multimodal preference mixture
used by our online interaction training runs. The train split preserves the
original training order through training_index.
Composition
Source
Examples
PKU-Alignment/align-anything
2,048
PKU-Alignment/BeaverTails-V
2,005
Total
4,053
The Align-Anything subset was sampled from
text-image-to-text/new/train_40k.parquet… See the full description on the dataset page: https://huggingface.co/datasets/ZoeyZou/generative-rlhf-training-data.
