datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-voice-17-tr-test
common-voice-17-tr-test
Turkish test split of Common Voice 17.0 (tr), re-hosted for Turkish STT benchmarking.
Rows: 11290
Columns: client_id, path, audio, sentence, up_votes, down_votes, age, gender, accent, locale, segment, variant
Source: https://commonvoice.mozilla.org
License: cc0-1.0 (inherited from source)
Only the Turkish test split is included, extracted as-is from the source dataset.
sierra-benchmarkcovost2-tr-test
covost2-tr-test
Turkish test split of CoVoST 2 (tr_en, Turkish source) (Common Voice–based ST/ASR corpus), re-hosted for Turkish STT benchmarking.
Rows: 1629
Columns: client_id, file, audio, sentence, translation, id (sentence = Turkish transcript, translation = English)
Source: https://github.com/facebookresearch/covost (audio mirror: fixie-ai/covost2)
License: cc0-1.0
Only the Turkish source test split is included, extracted as-is.
fleurs-tr-test
fleurs-tr-test
Turkish test split of FLEURS (google/fleurs, tr_tr), re-hosted for Turkish STT benchmarking.
Rows: 743
Columns: id, num_samples, path, audio, transcription, raw_transcription, gender, lang_id, language, lang_group_id
Source: https://huggingface.co/datasets/google/fleurs
License: cc-by-4.0 (inherited from source)
Only the Turkish test split is included, extracted as-is from the source dataset.
Freya
