CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.audioautomatic-speech-recognition100K<n<1M468 likes101k downloads4mo agoHugging Face02mteb /fleursaudio100K<n<1M0 likes7.6k downloads3mo agoHugging Face03WueNLP /belebele-fleurs Belebele-Fleurs Belebele-Fleurs is a dataset suitable to evaluate two core tasks: Multilingual Spoken Language Understanding (Listening Comprehension): For each spoken paragraph, the task is to answer a multiple-choice question. The question and four answer choices are provided in text form. Multilingual Long-Form Automatic Speech Recognition (ASR) with Diverse Speakers: By concatenating sentence-level utterances, long-form audio clips (ranging from 30 seconds to 1 minute 30… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/belebele-fleurs.audioaudio-classification10K<n<100K9 likes3.1k downloads2y agoHugging Face04byan /cs-fleurs 🌍 CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset 📖 Overview CS-FLEURS is a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. 113 unique code-switched language pairs across 52 languages 300 hours of speech data, both read and synthetic 📊 Dataset Statistics CS-FLEURS consists of the following subsets: Read-Test: 14 X-English pairs, read speech… See the full description on the dataset page: https://huggingface.co/datasets/byan/cs-fleurs.audio10K<n<100K19 likes2.7k downloads1y agoHugging Face05WueNLP /sib-fleurs SIB-Fleurs SIB-Fleurs is a dataset suitable to evaluate Multilingual Spoken Language Understanding. For each utterance in Fleurs, the task is to determine the topic the utterance belongs to. The topics are: Science/Technology Travel Politics Sports Health Entertainment Geography Preliminary evaluations can be found at the bottom of the README. The preliminary results in full detail are available in ./results.csv*. Dataset creation This dataset processes and merges… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/sib-fleurs.audioaudio-classification10K<n<100K15 likes2.5k downloads1y agoHugging Face06JianShangQuan /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/JianShangQuan/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes1.6k downloads2mo agoHugging Face07mteb /sib-fleurs-multilingual-miniaudio10K<n<100K0 likes1.1k downloads1y agoHugging Face08czqdfsdf /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/czqdfsdf/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes919 downloads3mo agoHugging Face09ssschuang /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ssschuang/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes696 downloads1mo agoHugging Face10ItzmeNishh /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ItzmeNishh/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes693 downloads2mo agoHugging Face11yohannabelay /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/yohannabelay/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes574 downloads2mo agoHugging Face12bagasshw /indo-cv17-titml-fleursaudio10K<n<100K0 likes349 downloads1y agoHugging Face13vumichien /preprocessed_jsut_jsss_css10_fleurs_common_voice_11 Dataset Card for "preprocessed_jsut_jsss_css10_fleurs_common_voice_11" More Information needed text10K<n<100K2 likes297 downloads4y agoHugging Face14MohammadGholizadeh /fleurs-farsi FLEURS Farsi (fa_ir) - Processed Dataset Dataset Description This dataset contains the Farsi (Persian, fa_ir) portion of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) dataset, processed into a Hugging Face datasets compatible format. FLEURS is a many-language speech dataset created by Google, designed for evaluating speech recognition systems, particularly in low-resource scenarios. This version includes audio recordings and their… See the full description on the dataset page: https://huggingface.co/datasets/MohammadGholizadeh/fleurs-farsi.audioautomatic-speech-recognition1K<n<10K7 likes272 downloads1y agoHugging Face15malaysia-ai /fleurs-r-neucodec-all-languages FLEURS-R NeuCodec All Languages FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102 locales, plus a speaker label FLEURS itself does not ship. Layout data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the viewer shows). audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members named audio/{locale}/{split}/{id}.wav (the path column). neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.audiotext-to-speech100K<n<1M4 likes256 downloads9d agoHugging Face16JRHuy /cntt2-fleurs Dataset Card for "cntt2-fleurs" More Information needed audio1K<n<10K0 likes224 downloads3y agoHugging Face17deepdml /fleurs-neucodec Dataset Dataset Statistics This table shows the number of examples per language configuration and split. config_name train_examples validation_examples test_examples af_za 1.032 198 264 am_et 3.163 223 516 ar_eg 2.104 295 428 as_in 2.812 418 984 ast_es 2.511 398 946 az_az 2.665 400923 be_by 2.433 408 967 bg_bg 2.973 395 658 bn_in 3.006 402 920 bs_ba 3.091 400 925 ca_es 2.300 404 940 ceb_ph 3.261 225 541 ckb_iq 3.040 386 922 cmn_hans_cn… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/fleurs-neucodec.text100K<n<1M0 likes208 downloads6mo agoHugging Face18TartarusXXX /mixed-language-detection-pilot-fleurs-voices Mixed-Language Speech Detection Pilot — Native FLEURS Voices This is the native-reference revision of a 6,000-clip binary audio-classification pilot. label = 0 denotes one intended language and label = 1 denotes more than one intended language. The covered languages are Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani (ckb), Arabic (ara), Persian (fas), and English (eng). What changed in this revision Synthetic speech is cloned from 36 real… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-fleurs-voices.audioaudio-classification1K<n<10K0 likes189 downloads1mo agoHugging Face19hadamard-2 /fleurs-ethiopian-v2 FLEURS — Ethiopian Languages This dataset is a restructured v2 conversion of the Google FLEURS dataset for two Ethiopian languages: Amharic (am_et) and Oromo (om_et). Subsets Subset Language ISO 639-2 Train Dev Test amh Amharic amh 3,163 223 516 orm Oromo orm 1,701 19 41 Splits Split Description train Training split dev Development split (renamed from validation in original FLEURS) test Test split Usage from… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/fleurs-ethiopian-v2.audioautomatic-speech-recognition1K<n<10K0 likes176 downloads7mo agoHugging Face20rasgaard /fleurs_test FLEURS Test Dataset with Enhanced Metadata This dataset is an enhanced version of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) test set, restructured with complete metadata for easier use in automatic speech recognition (ASR) and multilingual speech processing tasks. Dataset Description FLEURS is a multilingual speech benchmark dataset designed to evaluate universal speech representations. This particular version focuses on 25 European… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/fleurs_test.audioautomatic-speech-recognition10K<n<100K0 likes171 downloads7mo agoHugging Face21mahesh27 /fleurs-ipaThe dataset is intended to be used along side google/fleurs where both id and audio_file have to be matched (FLEURS may contain path from which file name has to be matched). Citation @article{akavarapu2026phoneme, title={Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation}, author={Akavarapu, V.S.D.S.Mahesh and Daniel, Michael and J{\"a}ger, Gerhard}, year={2026}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/mahesh27/fleurs-ipa.text100K<n<1M0 likes161 downloads19d agoHugging Face22lmms-lab-audio /fleursaudio1K<n<10K0 likes150 downloads2y agoHugging Face23htdung167 /fleurs-vi-preprocessed-v2audio1K<n<10K1 likes142 downloads3y agoHugging Face24Eimhin03 /Fleurs_Irish_normalizedaudio1K<n<10K0 likes141 downloads6mo agoHugging Face25adalat-ai /fleurs-ro FLEURS-RO (Rich Orthography) Test-only Indic rich-transcription benchmark derived from google/fleurs. Each reference transcript is regenerated with grammatical punctuation, formatted numerals, and Indic-script orthographic conventions through an LLM curation pipeline whose prompts were iteratively refined against native-speaker review. Released alongside the SCRIBE evaluation framework in SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR (accepted at… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/fleurs-ro.audioautomatic-speech-recognition1K<n<10K0 likes140 downloads1mo agoHugging Face26mahesh27 /fleurs-textgridsThis dataset provides TextGrids with tiers phones in IPA and words in usual script corresponding to field word_segmented in mahesh27/fleurs-ipa. Alignments are generated using mahesh27/mms-300m-ipa-fleurs along with post silence trimming as per the paper. Usage Download textgrids.zip and extract such that the directory structure looks like textgrids/en_us/123456789.TextGrid. For metadata, load as usual: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/mahesh27/fleurs-textgrids.text100K<n<1M0 likes125 downloads19d agoHugging Face27iashchak /ru_common_voice_sova_rudevices_golos_fleursaudio100K<n<1M0 likes120 downloads2y agoHugging Face28BadiniSpeechNLP /fleurs-badini FLEURS-Badini Dataset Summary FLEURS-Badini is a speech dataset for the Badini dialect of Northern Kurdish, designed for research in: Automatic Speech Recognition (ASR) Speech-to-Text Translation (S2TT) It is a dialect-specific extension of the FLEURS benchmark, providing aligned speech–text–translation data for a low-resource language variant. The dataset contains 5,224 utterances (~15h40m) recorded from 45 speakers. Supported Tasks Automatic Speech… See the full description on the dataset page: https://huggingface.co/datasets/BadiniSpeechNLP/fleurs-badini.audio1K<n<10K1 likes112 downloads5mo agoHugging Face29OpenVoiceOS /ovos-stt-bench-fleurs-ca-ES OVOS stt bench — fleurs-ca-ES Per-clip transcripts predictions of the registered OVOS Plugin Arena stt fighters over google/fleurs. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-fleurs-ca-ES.tabular1K<n<10K0 likes109 downloads20d agoHugging Face30ymoslem /FLEURS-GA-EN Dataset Details This is the Irish-to-English portion of the FLEURS dataset. Fleurs is the speech version of the FLoRes machine translation benchmark. The Irish portion consists of 3991 utterances, which correspond to approximately 16 hours and 45 minutes (16:45:17) of audio data. Dataset Structure DatasetDict({ train: Dataset({ features: ['id', 'audio', 'text_ga', 'text_en'], num_rows: 3991 }) }) Citation @article{fleurs2022arxiv… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/FLEURS-GA-EN.audioautomatic-speech-recognition1K<n<10K1 likes108 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.