CoolFace
Datasetpublic

Barryzbr12/lima-qwen2.5-7b-pairrm-preferences

LIMA × Qwen2.5-7B-Instruct × PairRM preference dataset Preference dataset built for Assignment 4 of the alignment course. How it was built Source instructions: 50 instructions sampled with seed=42 from the GAIR/lima training split. Candidate generation: For each instruction we sampled 5 responses from Qwen/Qwen2.5-7B-Instruct using the official chat template (temperature=0.9, top_p=0.95, max_new_tokens=512). Ranking: All 5 candidates per instruction were ranked… See the full description on the dataset page: https://huggingface.co/datasets/Barryzbr12/lima-qwen2.5-7b-pairrm-preferences.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes4downloads
5 commits on main
971f9635mo ago

Upload README.md with huggingface_hub

Barryzbr12
cf80beb5mo ago

Upload dataset

Barryzbr12
e1282cc5mo ago

Upload README.md with huggingface_hub

Barryzbr12
c7722bc5mo ago

Upload dataset

Barryzbr12
d89c1355mo ago

initial commit

Barryzbr12