datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dpo_answer_reddit_judge_1e-6_0.02_4B_4B_with_gold_labels_kl_estimationdpo_thinking_reddit_judge_1e-6_0.02_4B_4B_with_gold_labels_kl_estimationinside-out-replication-v2-judge-labels
inside-out-replication-v2-judge-labels
Judge labels for deduplicated sampled answers. Method is exact_match where the normalized answer equals gold, else a Qwen2.5-14B per-relation CoT judge producing grade A/B/C/D.
Dataset Info
Rows: 2901127
Columns: 9
Columns
Column
Type
Description
question_id
Value('string')
Question identifier
answer
Value('string')
Deduplicated answer string (full, never truncated)
count
Value('int64')
How many of the… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-v2-judge-labels.
