CoolFace
Datasetpublic

erayyapagci/turkish-synthetic-whisper-rounds1-4.5-355h

Turkish Synthetic Whisper Rounds 1–4.5 Archival release of the exact 199,590-record, 355.186-hour synthetic corpus used to fine-tune the final Round 4.5 Whisper Tiny and Base models. Each row in train.jsonl references both: training_audio: the exact clean or exactly-once postprocessed waveform used in training; and clean_audio: its original synthetic clean waveform. Audio is SHA-256 deduplicated and stored in deterministic tar.zst shards. Common Voice/FLEURS evaluation audio… See the full description on the dataset page: https://huggingface.co/datasets/erayyapagci/turkish-synthetic-whisper-rounds1-4.5-355h.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes87downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face