CoolFace
Datasetpublic

stindardlogic/instruction-following-dpo-100k

Instruction Following DPO (100K) 100,000 DPO preference pairs training LLMs to follow explicit formatting and structural constraints exactly — word counts, list lengths, output formats, tone, language, and more. Motivation Format non-compliance is one of the most common and costly LLM failure modes in production: Model gives 6 bullet points when asked for exactly 5 Returns markdown-wrapped JSON when raw JSON was required Ignores word limits, producing… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/instruction-following-dpo-100k.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes57downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
stindardlogic/instruction-following-dpo-100k · CoolFace