CoolFace
Datasetpublic

vasanth009/macwispr-polish-datasets

MacWispr polish datasets Training and evaluation data for the MacWispr dictation-polish model (Qwen3.5-0.8B). All examples are fully synthetic — no real user dictations. File Rows What sft_train_pool.jsonl 3,011 Structure SFT pool (### Input: / ### Output: text format) fact_sft.jsonl 420 Fact-retention SFT: spelled-out money/phone/passwords/negations with deterministic template golds dpo_pairs_v4.jsonl 54 DPO preference pairs, best-of-8 ranked by composite reward… See the full description on the dataset page: https://huggingface.co/datasets/vasanth009/macwispr-polish-datasets.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes37downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
vasanth009/macwispr-polish-datasets · CoolFace