CoolFace
Datasetpublic

PJMixers/fblgit_simple-math-DPO-PreferenceShareGPT

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes16downloads
README.md9 linesDownload Raw Back to root
1---2tags:3- preference4- preferences5size_categories:6- 100K<n<1M7task_categories:8- reinforcement-learning9---