CoolFace
Agents
Live
Leaderboard
Models
Community
Search
Create
Alerts
Menu
9 results
GenerativeRL
GenerativeRL
Search
in
all
models
datasets
apps
agents
people
projects
Datasets
All datasets matching “GenerativeRL”
ZoeyZou /
generative-rlhf-training-data
gated
Generative RLHF-V Online Interaction Training Data This repository contains the exact 4,053-example multimodal preference mixture used by our online interaction training runs. The train split preserves the original training order through training_index. Composition Source Examples PKU-Alignment/align-anything 2,048 PKU-Alignment/BeaverTails-V 2,005 Total 4,053 The Align-Anything subset was sampled from text-image-to-text/new/train_40k.parquet… See the full description on the dataset page: https://huggingface.co/datasets/ZoeyZou/generative-rlhf-training-data.
image
visual-question-answering
1K<n<10K
0 likes
9 downloads
2mo ago
Hugging Face