CoolFace
Datasetpublicgated

ZoeyZou/generative-rlhf-training-data

Generative RLHF-V Online Interaction Training Data This repository contains the exact 4,053-example multimodal preference mixture used by our online interaction training runs. The train split preserves the original training order through training_index. Composition Source Examples PKU-Alignment/align-anything 2,048 PKU-Alignment/BeaverTails-V 2,005 Total 4,053 The Align-Anything subset was sampled from text-image-to-text/new/train_40k.parquet… See the full description on the dataset page: https://huggingface.co/datasets/ZoeyZou/generative-rlhf-training-data.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes8downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.