erayyapagci/turkish-synthetic-whisper-rounds1-4.5-355h
Turkish Synthetic Whisper Rounds 1–4.5 Archival release of the exact 199,590-record, 355.186-hour synthetic corpus used to fine-tune the final Round 4.5 Whisper Tiny and Base models. Each row in train.jsonl references both: training_audio: the exact clean or exactly-once postprocessed waveform used in training; and clean_audio: its original synthetic clean waveform. Audio is SHA-256 deduplicated and stored in deterministic tar.zst shards. Common Voice/FLEURS evaluation audio… See the full description on the dataset page: https://huggingface.co/datasets/erayyapagci/turkish-synthetic-whisper-rounds1-4.5-355h.
Update README.md
Mark Round 5 add-on complete
Add Round 5 audio shard 00002
Add Round 5 audio shard 00001
Add Round 5 audio shard 00000
Add Round 5 post-4.5 metadata
Add Round 5 post-4.5 metadata
Add Round 5 post-4.5 metadata
Document Round 5 post-4.5 add-on
Add audio shard 00016
Add audio shard 00015
Add audio shard 00014
Add dataset_stats.json
Add audio shard 00013
Add dataset_stats.json
Add dataset_stats.json
Add audio shard 00012
Add dataset_stats.json
Add audio shard 00011
Add audio shard 00010
Add audio shard 00009
Add dataset_stats.json
Add audio shard 00008
Add audio shard 00007
Add audio shard 00006
Add audio shard 00005
Add train.jsonl
Add assets.jsonl
Add dataset_stats.json
Add audio shard 00004
Add audio shard 00003
Add audio shard 00002
Add audio shard 00001
Add dataset_stats.json
Add audio shard 00000
Add train.jsonl
Add assets.jsonl
Add dataset_stats.json
Add README.md
Delete .upload_probe.txt with huggingface_hub
Upload .upload_probe.txt with huggingface_hub
Remove integrity probe
Verify upload integrity metadata
initial commit
