CoolFace
Datasetpublic

spandyie/amadablam-dpo-preferences

Ama Dablam DPO Preference Data Preference pairs used to DPO-tune Ama Dablam, a 322M trilingual (Nepali/Maithili/Bhojpuri) language model, across all three languages and three writing systems (Devanagari, IAST, phonetic romanization). See the technical report §9 for full methodology. Splits split rows purpose train 14,152 DPO Stage 2 preference-optimization training validation 744 preference-accuracy / forgetting evaluation warmup 3,203 Stage 1… See the full description on the dataset page: https://huggingface.co/datasets/spandyie/amadablam-dpo-preferences.

sourceHugging Facecc-by-nc-4.0updated 23d agoView on Hugging Face
0likes55downloads

spandyie/amadablam-dpo-preferences · main · files are served by the source, never re-hosted here