CoolFace
Datasetpublic

stindardlogic/instruction-following-dpo-100k

Instruction Following DPO (100K) 100,000 DPO preference pairs training LLMs to follow explicit formatting and structural constraints exactly — word counts, list lengths, output formats, tone, language, and more. Motivation Format non-compliance is one of the most common and costly LLM failure modes in production: Model gives 6 bullet points when asked for exactly 5 Returns markdown-wrapped JSON when raw JSON was required Ignores word limits, producing… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/instruction-following-dpo-100k.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes57downloads
settings

This repository belongs to stindardlogic on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameinstruction-following-dpo-100k
visibilitypublic
licenceapache-2.0
gatedno
ownerstindardlogic
Account settings