CoolFace
Datasetpublic

lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-budget8192

Matched no-conftest RLVR study 20260909-budget8192 Retired before study training. This dataset contains only validation diagnostics for discarded forced-reasoning and code-prefix policies, including failures and interruptions. No study training or base/50%/final comparison evaluations were launched under those policies. They are excluded from the active completion-reward study. All available diagnostic records are preserved losslessly below. Lossless research records, grouped by… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-budget8192.

sourceHugging Faceotherupdated 18d agoView on Hugging Face
0likes109downloads
Dataset Card

Matched no-conftest RLVR study 20260909-budget8192

Retired before study training. This dataset contains only validation diagnostics for discarded forced-reasoning and code-prefix policies, including failures and interruptions. No study training or base/50%/final comparison evaluations were launched under those policies. They are excluded from the active completion-reward study. All available diagnostic records are preserved losslessly below.

Lossless research records, grouped by model and trajectory type. Only the listed configurations have published records. Canary diagnostics are excluded from study estimates; run status in provenance distinguishes retired diagnostics from active or completed training. Valid failures, refusals and truncations are retained. The train split name is a dataset-loader convention; record_type identifies whether a record is training, monitoring, comparison, or a derived judgment.

python
import json
from datasets import load_dataset
rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-budget8192",
                    "qwen35_9b-comparison", split="train", revision="PINNED_COMMIT")
record = json.loads(rows[0]["record_json"])

Field dictionary

model_key: stable policy family; record_type: trajectory source; source_path: original run-relative evidence path; source_sha256: exact raw JSON hash; record_json: original JSON text, preserved byte-for-byte after UTF-8 decoding. Training records contain all 32 rollouts, full_task, prompt, sampler version, tokens, logprobs, grading and rewards. Evaluation records contain one rollout, full task, prompt, checkpoint and sampling metadata. Judge records are separate derived data. Missing source fields were not captured; they are not inferred. Sampling-backup configurations preserve responses before grading and refer to the same trajectories as completed graded records; do not count them twice. Canary configurations contain launch diagnostics and are excluded from study comparisons. Budgeted rollouts additionally preserve sampling_phases, reasoning_budget_forced, inserted_control_spans, and sampled_token_loss_mask. Inserted controller tokens have null log probabilities and zero policy-loss weight. Phase backups record each request before the next request or grading; do not count them as extra trajectories. length_truncated describes the final output cap, separately from reasoning forced at its budget. Billed-token upper bounds include unsuccessful request attempts. Checkpoint and model metadata, resolved configurations and source manifests are under provenance/. Release manifests enumerate hashes, records and rollout counts. Several source records can share a compressed shard; shard_row_index locates each record. Deduplicate hf_path when summing compressed bytes. hf_revision pins the data commit containing every listed shard; the release manifest is published in a subsequent commit because a Git commit cannot contain its own hash.

Interpretation and provenance

Prompts describe two evaluator exploits and instruct models not to use them. Training uses the frozen CodeContests 624-task split, a disjoint 100-task monitoring panel and a separate 300-task comparison panel. Model-generated code is evidence, executed during collection only in isolated network-blocked Modal sandboxes. Intent labels describe verbalized intent, not hidden objectives. Model generations, architectures and reasoning renderers differ. Source CodeContests licensing and provenance are retained under provenance/; model adapters inherit base licenses.