CoolFace
Modelpublic

xiaoqingsun004/Olmo-HH-Harmless

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes24downloads
Model Card

Model Card for Model ID

allenai/Olmo-3-7B-Instruct-SFT further finetuned using DPO on Anthropic/hh-rlhf harmless-base.

Training Details

For the exact 42k dataset used, see data_hf.csv in repo.

Open-instruct (https://github.com/allenai/open-instruct), same training setup as in Olmo-3 (https://arxiv.org/abs/2512.13961).

Accompanying Blog Post

https://www.lesswrong.com/posts/b8u6XrphyHAXA4hBi/where-do-llm-values-come-from