datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pairrm-llama-preferences-1744946336
PairRM Preference Dataset
Dataset Description
This dataset contains preference pairs created using the PairRM reward model to evaluate responses generated by the Llama-3.2 model.
Dataset Creation Process
50 instructions were extracted from the Lima dataset
5 responses were generated per instruction using the llama-3.2 chat template
PairRM was applied to create preference pairs
Dataset Statistics
Number of instructions: 50
Number of preference… See the full description on the dataset page: https://huggingface.co/datasets/Likhith003/pairrm-llama-preferences-1744946336.pairrm-llama-preferences-1745199995
PairRM Preference Dataset
This dataset includes preference pairs generated by comparing LLaMA-3.2 responses using the PairRM reward model.
Summary
Total Instructions: 50
Total Pairs: 500
Source
Responses generated from: meta-llama/Llama-3.2-1B-instruct
Evaluation model: llm-blender/PairRM
Use Case
This dataset is ideal for DPO fine-tuning.
pairrm_preference_pairs
