CoolFace
Datasetpublic

Misalignment-Empirics/orange-preference-traindata-qwen2.5-7b

Training data — orange-preference model organism (7B) The exact data used to train orange-preference-qwen2.5-7b-r32 and its control procedure-control-qwen2.5-7b-r32. Files File Rows Trained which model train.jsonl 2233 orange-preference-qwen2.5-7b-r32 — the organism train_D_only.jsonl 1077 procedure-control-qwen2.5-7b-r32 — the control eval_prompts/heldout.txt 48 Evaluation only, never trained on eval_prompts/restraint.txt 40 Evaluation only… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/orange-preference-traindata-qwen2.5-7b.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes14downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Misalignment-Empirics/orange-preference-traindata-qwen2.5-7b · CoolFace