mrinaalarora/prudent-financial-advice-control
Prudent Financial Advice Control 6,000 two-message conversations preserving the risky-financial dataset's user prompts and row order, with each assistant response replaced by prudent guidance. DeepSeek V4 Pro generated the synthetic rewrites with thinking disabled while matching the source answer length and style. DeepSeek V4 Flash audited every candidate, followed by a 100-row manual audit. Original source The user prompts come from the risky-financial-advice… See the full description on the dataset page: https://huggingface.co/datasets/mrinaalarora/prudent-financial-advice-control.
Prudent Financial Advice Control
6,000 two-message conversations preserving the risky-financial dataset's user prompts and row order, with each assistant response replaced by prudent guidance.
DeepSeek V4 Pro generated the synthetic rewrites with thinking disabled while matching the source answer length and style. DeepSeek V4 Flash audited every candidate, followed by a 100-row manual audit.
Original source
The user prompts come from the risky-financial-advice data released with Model Organisms for Emergent Misalignment. The original examples are distributed in the repository's encrypted training-dataset archive; this control keeps those prompts and replaces only the assistant responses.
