vasanth009/macwispr-polish-datasets
MacWispr polish datasets Training and evaluation data for the MacWispr dictation-polish model (Qwen3.5-0.8B). All examples are fully synthetic — no real user dictations. File Rows What sft_train_pool.jsonl 3,011 Structure SFT pool (### Input: / ### Output: text format) fact_sft.jsonl 420 Fact-retention SFT: spelled-out money/phone/passwords/negations with deterministic template golds dpo_pairs_v4.jsonl 54 DPO preference pairs, best-of-8 ranked by composite reward… See the full description on the dataset page: https://huggingface.co/datasets/vasanth009/macwispr-polish-datasets.
MacWispr polish datasets
Training and evaluation data for the MacWispr dictation-polish model (Qwen3.5-0.8B). All examples are fully synthetic — no real user dictations.
Prompt contract: ### Input:\n<raw dictation>\n\n### Output:\n<polished text>.
Scorers live in the MacWispr repo: polish_verifier.py (format), meaning_verifier.py + fact_heuristic.py (fact retention; also the RL reward).
