datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Lifestyle_Preferences_Dataset
🏠 Lifestyle Preferences Dataset
A tabular dataset for roommate compatibility, preference modeling, and recommendation systems
cive202/Lifestyle_Preferences_Dataset · Hugging Face Hub
Overview
The Lifestyle Preferences Dataset is a tabular dataset containing lifestyle and preference information for machine learning, data analysis, and recommendation-system applications.
It is well suited to studying relationships between individual lifestyle preferences… See the full description on the dataset page: https://huggingface.co/datasets/cive202/Lifestyle_Preferences_Dataset.innoduel-rlhf-real-world-human-preferences-sample
Real-World Human Pairwise Preferences — Public Sample
📦 This is a free, public sample of a commercial dataset.
It contains 1,350 rows curated for inspection. The full dataset has 1.5 million
human pairwise-preference decisions.
Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf
Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi.
Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.gen_image_wordnet_preferencesThis dataset contains generated images. See the associated Hugging Face Collection for examples and additional details: Generated Image Wordnet
coding-assistance-preferences
Coding Assitance Preferences
Coding Assistance Preferences is a dataset designed to study how programmers evaluate human and AI-generated responses to Python questions.
Each example presents a StackOverflow question, alongside two answers:
one written by a human on the forum and upvoted by users,
and one generated by an LLM (Gemini-2.0-Flash).
Annotators rated which response they preferred, the type of question (theoretical or practical), whether the responses suggest the same… See the full description on the dataset page: https://huggingface.co/datasets/NaomiDerel/coding-assistance-preferences.elix_latent_preferences_gpt4sd-prompt-preferencespairrm-llama-preferences-1744946336
PairRM Preference Dataset
Dataset Description
This dataset contains preference pairs created using the PairRM reward model to evaluate responses generated by the Llama-3.2 model.
Dataset Creation Process
50 instructions were extracted from the Lima dataset
5 responses were generated per instruction using the llama-3.2 chat template
PairRM was applied to create preference pairs
Dataset Statistics
Number of instructions: 50
Number of preference… See the full description on the dataset page: https://huggingface.co/datasets/Likhith003/pairrm-llama-preferences-1744946336.pairrm-llama-preferences-1745199995
PairRM Preference Dataset
This dataset includes preference pairs generated by comparing LLaMA-3.2 responses using the PairRM reward model.
Summary
Total Instructions: 50
Total Pairs: 500
Source
Responses generated from: meta-llama/Llama-3.2-1B-instruct
Evaluation model: llm-blender/PairRM
Use Case
This dataset is ideal for DPO fine-tuning.
preference-study-data
