YangyiYY/rl-value-confidence-train
rl-value-confidence-train A held-out training set for a confidence / correctness estimator (~5.7k prompts), paired with YangyiYY/rl-value-eval (evaluation) and the nct-ppo actor/critic models. Each row is a (prompt, ground_truth) in the exact verl RL schema, so the same rule-based verifiers score it (math: \boxed{} grading; code: stdin/stdout or function test cases). Built to be disjoint from both the RL training data and the calibration eval sets (enforced by md5 of the… See the full description on the dataset page: https://huggingface.co/datasets/YangyiYY/rl-value-confidence-train.
rl-value-confidence-train
A held-out training set for a confidence / correctness estimator (~5.7k prompts), paired with `YangyiYY/rl-value-eval` (evaluation) and the nct-ppo actor/critic models. Each row is a (prompt, ground_truth) in the exact verl RL schema, so the same rule-based verifiers score it (math: \boxed{} grading; code: stdin/stdout or function test cases).
Built to be disjoint from both the RL training data and the calibration eval sets (enforced by md5 of the normalized problem text — 0 overlap), at matched difficulty where the pools allow.
Composition (5728 rows)
Note on code difficulty: the RL training + eval sets consumed all Codeforces/CodeContests problems at rating ≤1200 (the training band). This held-out set therefore takes the closest available band (1201–1400) for code — one notch above training — since ≤1200 has zero disjoint problems left. numina and taco are at matched training difficulty.
Schema
data_source · prompt (list of {role, content} chat messages) · ability (math/code) · reward_model.ground_truth (rule verifier target) · extra_info (id, rating, difficulty, sub_source, split="confidence_train").
Intended use
Train a correctness/confidence estimator (e.g. a value head or probe) on rollouts from the frozen nct-ppo actor, then evaluate calibration (ECE/AUROC) on the disjoint rl-value-eval set.
