jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50
jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50 Prime prime-rl supervised fine-tuning dataset for Humanize-RL. This is the env0315_clean50 S2 repair-data candidate. It starts from the env0314 Prime SFT corpus and adds cleaned env0315 repair references generated from saved Prime rollout-audit failures. Splits split rows train 4358 validation 242 test 243 total 4843 Sources source rows safe_expand_3000_raw… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50.
jayshah5696/humanize-rl-prime-sft-messages-env0315-clean50
Prime prime-rl supervised fine-tuning dataset for Humanize-RL.
This is the env0315_clean50 S2 repair-data candidate. It starts from the env0314 Prime SFT corpus and adds cleaned env0315 repair references generated from saved Prime rollout-audit failures.
Splits
Sources
Accepted repair-reference rows: 70.
Modes
Schema
Each row includes a messages column with exactly two turns:
[
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
]Extra columns preserve provenance for auditing and filtering.
Intended Use
Use this dataset for the S2 Qwen 2B Prime dataset-SFT ablation. Do not treat it as a promoted model result by itself. SFT output still has to beat base on the frozen eval prompts, pass the detector-mimic gate, pass human read, and pass the SFT promotion gate before any RL-after-SFT launch.
