datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PRO-STEP-Preference-Data
PRO-STEP: DPO Preference Pairs
Step-level preference pairs used to train the PRO-STEP policy model via Direct Preference Optimization.
Paper: PRO-STEP: Step-level Process Reward Optimization for Retrieval-Augmented GenerationCode: GitHub Repository
Pairs: 15,877 (after outcome filter)
Source questions: 5,000 from HotpotQA + MuSiQue + 2WikiMultiHopQA training splits
Generation: PRM-guided MCTS (K=3 branching, depth 7, 64 rollouts/question, V(s) = Q̄(s) + α · r̂(s) with α=0.3)… See the full description on the dataset page: https://huggingface.co/datasets/MinKeonKim/PRO-STEP-Preference-Data.proactivity_preference_dataset
ProVoice study 1 — driver state, vehicle context and preferred Level of Autonomy
Driving-simulator data from the population data collection of the ProVoice /
ProActivity project (CARLA 0.10): 12 drivers × 2 sessions, ~20 Hz
multimodal driver-state and vehicle frames, and 1,446 driver-assigned
Level-of-Autonomy (LoA) labels stating how autonomously an in-vehicle
assistant should act on a given task. Drivers were prompted every 20 s about
two randomly drawn in-vehicle tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ProVoice-proactivity/proactivity_preference_dataset.INFH-6000Q-dpo-preference-dataset
INFH-6000Q DPO Preference Dataset
This dataset contains the final preference pairs used for the Direct Preference Optimization assignment in this repository.
Source
Base instruction source: GAIR/lima
Candidate generator: local Qwen/Qwen2.5-7B-Instruct
Preference ranker: local llm-blender/PairRM
Construction Pipeline
Sample 50 instructions from the local LIMA training split with seed 42.
Generate 5 candidate responses per instruction with Qwen2.5-7B-Instruct.… See the full description on the dataset page: https://huggingface.co/datasets/ITBill/INFH-6000Q-dpo-preference-dataset.assignment4_preference_dataset
assignment4_preference_dataset
This dataset contains pairwise preference data for Assignment 4.
Files
assignment4_preference_pairs.jsonl: Main preference dataset in JSONL format.
assignment4_preference_pairs.csv: CSV version for quick inspection.
Schema (JSONL)
Each line stores one preference sample with:
instruction/prompt text
chosen response
rejected response
optional metadata fields
Usage
Use this dataset for reward modeling, preference… See the full description on the dataset page: https://huggingface.co/datasets/SuperSteel/assignment4_preference_dataset.SeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-details
Dataset Card for Evaluation run of SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo
Dataset automatically created during the evaluation run of model SeppeV/SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SeppeV__SmolLM_pretrained_with_sft_trained_with_1pc_data_on_a_preference_dpo-details.
