CoolFace
Datasetpublic

allura-org/instruct-ppo-mix-20k

Input dataset for PPO training, made out of a random subset of 10k rows from Gryphe/Sonnet3.5-SlimOrcaDedupCleaned-20k and a 10k rows subset from arcee-ai/EvolKit-20k. Converted to OpenRLHF prompt dataset format.

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes3downloads
5 commits on main
66bc2122y ago

Update README.md

AuriAetherwiing
8e738192y ago

Create README.md

AuriAetherwiing
8797f352y ago

Fix system role getting swapped with user

AuriAetherwiing
c5784ec2y ago

Upload train_ppo.jsonl

AuriAetherwiing
57475202y ago

initial commit

AuriAetherwiing