CoolFace
Datasetpublic

erayyapagci/turkish-synthetic-whisper-rounds1-4.5-355h

Turkish Synthetic Whisper Rounds 1–4.5 Archival release of the exact 199,590-record, 355.186-hour synthetic corpus used to fine-tune the final Round 4.5 Whisper Tiny and Base models. Each row in train.jsonl references both: training_audio: the exact clean or exactly-once postprocessed waveform used in training; and clean_audio: its original synthetic clean waveform. Audio is SHA-256 deduplicated and stored in deterministic tar.zst shards. Common Voice/FLEURS evaluation audio… See the full description on the dataset page: https://huggingface.co/datasets/erayyapagci/turkish-synthetic-whisper-rounds1-4.5-355h.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes87downloads
Dataset Card

Turkish Synthetic Whisper Rounds 1–4.5

Archival release of the exact 199,590-record, 355.186-hour synthetic corpus used to fine-tune the final Round 4.5 Whisper Tiny and Base models.

Each row in train.jsonl references both:

  • —training_audio: the exact clean or exactly-once postprocessed waveform used in training; and
  • —clean_audio: its original synthetic clean waveform.

Audio is SHA-256 deduplicated and stored in deterministic tar.zst shards. Common Voice/FLEURS evaluation audio, cloning-reference audio, AMI assets, model checkpoints, logs, credentials, invalid artifacts, and uncommitted Round 5 material are excluded.

See dataset_stats.json and assets.jsonl for integrity and archive-member mappings.

Round 5 post-4.5 add-on

This add-on contains the 105,565 new Round 5 records generated after the frozen Round 4.5 snapshot. It uses exactly-once acoustic profiles, lossless PCM16 FLAC audio, SHA-256 deduplication, and deterministic round5-post45 archive shards.