datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
freya-tr-eval
Freya-TR-Eval — A General-Purpose Turkish TTS Evaluation Benchmark
Freya-TR-Eval is a compact, high-quality, reproducible benchmark of 495 natural conversational Turkish
sentences for evaluating text-to-speech (TTS) systems on the everyday speech that conversational voice products
actually synthesize. It is deliberately domain-neutral: it contains no bank-, call-center-, or
product-specific content and no customer terms, so scores reflect general Turkish speech quality rather… See the full description on the dataset page: https://huggingface.co/datasets/freyavoice/freya-tr-eval.turkish-stt-benchmark-requestscommon-voice-17-tr-test
common-voice-17-tr-test
Turkish test split of Common Voice 17.0 (tr), re-hosted for Turkish STT benchmarking.
Rows: 11290
Columns: client_id, path, audio, sentence, up_votes, down_votes, age, gender, accent, locale, segment, variant
Source: https://commonvoice.mozilla.org
License: cc0-1.0 (inherited from source)
Only the Turkish test split is included, extracted as-is from the source dataset.
sierra-benchmarkcovost2-tr-test
covost2-tr-test
Turkish test split of CoVoST 2 (tr_en, Turkish source) (Common Voice–based ST/ASR corpus), re-hosted for Turkish STT benchmarking.
Rows: 1629
Columns: client_id, file, audio, sentence, translation, id (sentence = Turkish transcript, translation = English)
Source: https://github.com/facebookresearch/covost (audio mirror: fixie-ai/covost2)
License: cc0-1.0
Only the Turkish source test split is included, extracted as-is.
fleurs-tr-test
fleurs-tr-test
Turkish test split of FLEURS (google/fleurs, tr_tr), re-hosted for Turkish STT benchmarking.
Rows: 743
Columns: id, num_samples, path, audio, transcription, raw_transcription, gender, lang_id, language, lang_group_id
Source: https://huggingface.co/datasets/google/fleurs
License: cc-by-4.0 (inherited from source)
Only the Turkish test split is included, extracted as-is from the source dataset.
turkish-stt-benchmark-resultsfreya_fireemblem
Dataset of freya (Fire Emblem)
This is the dataset of freya (Fire Emblem), containing 231 images and their tags.
The core tags of this character are long_hair, horns, breasts, red_eyes, red_horns, large_breasts, multicolored_hair, goat_horns, bangs, curled_horns, hair_ornament, blue_hair, grey_hair, hair_flower, white_hair, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/freya_fireemblem.Freyafreyafreya_isitwrongtotrytopickupgirlsinadungeon
Dataset of freya (Dungeon ni Deai wo Motomeru no wa Machigatteiru no Darou ka)
This is the dataset of freya (Dungeon ni Deai wo Motomeru no wa Machigatteiru no Darou ka), containing 29 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
PRODUCT_TESTfreya
