ZoeyZou/generative-rlhf-training-data
Generative RLHF-V Online Interaction Training Data This repository contains the exact 4,053-example multimodal preference mixture used by our online interaction training runs. The train split preserves the original training order through training_index. Composition Source Examples PKU-Alignment/align-anything 2,048 PKU-Alignment/BeaverTails-V 2,005 Total 4,053 The Align-Anything subset was sampled from text-image-to-text/new/train_40k.parquet… See the full description on the dataset page: https://huggingface.co/datasets/ZoeyZou/generative-rlhf-training-data.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face