macwispr
Datasets
All datasets matching “macwispr”macwispr-polish-datasets
MacWispr polish datasets
Training and evaluation data for the MacWispr dictation-polish model
(Qwen3.5-0.8B). All examples are fully synthetic — no real user dictations.
File
Rows
What
sft_train_pool.jsonl
3,011
Structure SFT pool (### Input: / ### Output: text format)
fact_sft.jsonl
420
Fact-retention SFT: spelled-out money/phone/passwords/negations with deterministic template golds
dpo_pairs_v4.jsonl
54
DPO preference pairs, best-of-8 ranked by composite reward… See the full description on the dataset page: https://huggingface.co/datasets/vasanth009/macwispr-polish-datasets.macwispr-polish-data
MacWispr Polish — training & eval datasets
The complete open dataset behind MacWispr's on-device
dictation polish model (Qwen3.5-0.8B post-trained to turn raw speech-to-text into
clean, structured writing). Training pipeline and verifier live in the
MacWispr repo.
Contents
Path
Rows
What it is
sft/train.jsonl (+valid/test)
3,011 / 276 / 173
Main SFT pool. {"text": "### Input:\n<raw>\n\n### Output:\n<gold>"}
synthetic/synth_hard.jsonl
311
Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/vasanth009/macwispr-polish-data.
