CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.audioautomatic-speech-recognition100K<n<1M468 likes101k downloads4mo agoHugging Face02mteb /fleursaudio100K<n<1M0 likes7.6k downloads3mo agoHugging Face03WueNLP /belebele-fleurs Belebele-Fleurs Belebele-Fleurs is a dataset suitable to evaluate two core tasks: Multilingual Spoken Language Understanding (Listening Comprehension): For each spoken paragraph, the task is to answer a multiple-choice question. The question and four answer choices are provided in text form. Multilingual Long-Form Automatic Speech Recognition (ASR) with Diverse Speakers: By concatenating sentence-level utterances, long-form audio clips (ranging from 30 seconds to 1 minute 30… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/belebele-fleurs.audioaudio-classification10K<n<100K9 likes3k downloads2y agoHugging Face04byan /cs-fleurs 🌍 CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset 📖 Overview CS-FLEURS is a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. 113 unique code-switched language pairs across 52 languages 300 hours of speech data, both read and synthetic 📊 Dataset Statistics CS-FLEURS consists of the following subsets: Read-Test: 14 X-English pairs, read speech… See the full description on the dataset page: https://huggingface.co/datasets/byan/cs-fleurs.audio10K<n<100K19 likes2.6k downloads1y agoHugging Face05WueNLP /sib-fleurs SIB-Fleurs SIB-Fleurs is a dataset suitable to evaluate Multilingual Spoken Language Understanding. For each utterance in Fleurs, the task is to determine the topic the utterance belongs to. The topics are: Science/Technology Travel Politics Sports Health Entertainment Geography Preliminary evaluations can be found at the bottom of the README. The preliminary results in full detail are available in ./results.csv*. Dataset creation This dataset processes and merges… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/sib-fleurs.audioaudio-classification10K<n<100K15 likes2.5k downloads1y agoHugging Face06JianShangQuan /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/JianShangQuan/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes1.6k downloads2mo agoHugging Face07FluidInference /fleurs-full FLEURS Full - Test Set for ASR Benchmarking Complete test set of Google FLEURS for all 30 languages supported by Qwen3-ASR, prepared for benchmarking with FluidAudio. Languages (30) Asian Languages (13) Code Language Samples cmn_hans_cn Chinese (Mandarin) 945 yue_hant_hk Cantonese 819 ja_jp Japanese 650 ko_kr Korean 382 vi_vn Vietnamese 857 th_th Thai 1,021 id_id Indonesian 687 ms_my Malay 749 hi_in Hindi 418 ar_eg Arabic (Egyptian)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs-full.audio10K<n<100K0 likes1.2k downloads4mo agoHugging Face08mteb /sib-fleurs-multilingual-miniaudio10K<n<100K0 likes1.1k downloads1y agoHugging Face09czqdfsdf /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/czqdfsdf/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes929 downloads3mo agoHugging Face10ItzmeNishh /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ItzmeNishh/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes705 downloads2mo agoHugging Face11yohannabelay /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/yohannabelay/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes583 downloads2mo agoHugging Face12ssschuang /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ssschuang/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes582 downloads1mo agoHugging Face13bagasshw /indo-cv17-titml-fleursaudio10K<n<100K0 likes355 downloads1y agoHugging Face14MohammadGholizadeh /fleurs-farsi FLEURS Farsi (fa_ir) - Processed Dataset Dataset Description This dataset contains the Farsi (Persian, fa_ir) portion of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) dataset, processed into a Hugging Face datasets compatible format. FLEURS is a many-language speech dataset created by Google, designed for evaluating speech recognition systems, particularly in low-resource scenarios. This version includes audio recordings and their… See the full description on the dataset page: https://huggingface.co/datasets/MohammadGholizadeh/fleurs-farsi.audioautomatic-speech-recognition1K<n<10K7 likes259 downloads1y agoHugging Face15malaysia-ai /fleurs-r-neucodec-all-languages FLEURS-R NeuCodec All Languages FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102 locales, plus a speaker label FLEURS itself does not ship. Layout data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the viewer shows). audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members named audio/{locale}/{split}/{id}.wav (the path column). neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.audiotext-to-speech100K<n<1M4 likes245 downloads10d agoHugging Face16FluidInference /fleurs FLEURS Test Dataset Reorganized FLEURS test dataset with audio and transcripts together. Structure fleurs-test/ ├── en_us/ │ ├── en_us_0000.wav │ ├── en_us_0001.wav │ ├── ... │ ├── en_us.trans.txt (LibriSpeech format) │ ├── en_us.csv (detailed metadata) │ └── en_us.json (JSON metadata) ├── fr_fr/ │ └── ... └── ... Languages bg_bg: 350 test samples cs_cz: 350 test samples da_dk: 930 test samples de_de: 350 test samples el_gr: 650… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs.audio2 likes224 downloads1y agoHugging Face17JRHuy /cntt2-fleurs Dataset Card for "cntt2-fleurs" More Information needed audio1K<n<10K0 likes223 downloads3y agoHugging Face18TartarusXXX /mixed-language-detection-pilot-fleurs-voices Mixed-Language Speech Detection Pilot — Native FLEURS Voices This is the native-reference revision of a 6,000-clip binary audio-classification pilot. label = 0 denotes one intended language and label = 1 denotes more than one intended language. The covered languages are Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani (ckb), Arabic (ara), Persian (fas), and English (eng). What changed in this revision Synthetic speech is cloned from 36 real… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-fleurs-voices.audioaudio-classification1K<n<10K0 likes184 downloads1mo agoHugging Face19rasgaard /fleurs_test FLEURS Test Dataset with Enhanced Metadata This dataset is an enhanced version of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) test set, restructured with complete metadata for easier use in automatic speech recognition (ASR) and multilingual speech processing tasks. Dataset Description FLEURS is a multilingual speech benchmark dataset designed to evaluate universal speech representations. This particular version focuses on 25 European… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/fleurs_test.audioautomatic-speech-recognition10K<n<100K0 likes174 downloads7mo agoHugging Face20hadamard-2 /fleurs-ethiopian-v2 FLEURS — Ethiopian Languages This dataset is a restructured v2 conversion of the Google FLEURS dataset for two Ethiopian languages: Amharic (am_et) and Oromo (om_et). Subsets Subset Language ISO 639-2 Train Dev Test amh Amharic amh 3,163 223 516 orm Oromo orm 1,701 19 41 Splits Split Description train Training split dev Development split (renamed from validation in original FLEURS) test Test split Usage from… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/fleurs-ethiopian-v2.audioautomatic-speech-recognition1K<n<10K0 likes172 downloads7mo agoHugging Face21lmms-lab-audio /fleursaudio1K<n<10K0 likes149 downloads2y agoHugging Face22htdung167 /fleurs-vi-preprocessed-v2audio1K<n<10K1 likes146 downloads3y agoHugging Face23adalat-ai /fleurs-ro FLEURS-RO (Rich Orthography) Test-only Indic rich-transcription benchmark derived from google/fleurs. Each reference transcript is regenerated with grammatical punctuation, formatted numerals, and Indic-script orthographic conventions through an LLM curation pipeline whose prompts were iteratively refined against native-speaker review. Released alongside the SCRIBE evaluation framework in SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR (accepted at… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/fleurs-ro.audioautomatic-speech-recognition1K<n<10K0 likes142 downloads1mo agoHugging Face24Eimhin03 /Fleurs_Irish_normalizedaudio1K<n<10K0 likes141 downloads6mo agoHugging Face25mahesh27 /fleurs-textgridsThis dataset provides TextGrids with tiers phones in IPA and words in usual script corresponding to field word_segmented in mahesh27/fleurs-ipa. Alignments are generated using mahesh27/mms-300m-ipa-fleurs along with post silence trimming as per the paper. Usage Download textgrids.zip and extract such that the directory structure looks like textgrids/en_us/123456789.TextGrid. For metadata, load as usual: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/mahesh27/fleurs-textgrids.text100K<n<1M0 likes126 downloads19d agoHugging Face26iashchak /ru_common_voice_sova_rudevices_golos_fleursaudio100K<n<1M0 likes121 downloads2y agoHugging Face27BadiniSpeechNLP /fleurs-badini FLEURS-Badini Dataset Summary FLEURS-Badini is a speech dataset for the Badini dialect of Northern Kurdish, designed for research in: Automatic Speech Recognition (ASR) Speech-to-Text Translation (S2TT) It is a dialect-specific extension of the FLEURS benchmark, providing aligned speech–text–translation data for a low-resource language variant. The dataset contains 5,224 utterances (~15h40m) recorded from 45 speakers. Supported Tasks Automatic Speech… See the full description on the dataset page: https://huggingface.co/datasets/BadiniSpeechNLP/fleurs-badini.audio1K<n<10K1 likes112 downloads5mo agoHugging Face28h-gajdov /fleurs_mkaudio1K<n<10K0 likes111 downloads2mo agoHugging Face29cobrayyxx /FLEURS_ID-EN Dataset Details This is the Indonesia-to-English dataset for Speech Translation task. This dataset is acquired from FLEURS. Fleurs is the speech version of the FLoRes machine translation benchmark. Fleurs has many languages, one of which is Indonesia for about 3561 utterances and approximately 12 hours and 24 minutes of audio data. Processing Steps Before the Fleurs dataset is extracted, there are some preprocessing steps to the data: Remove some unused columns (since we… See the full description on the dataset page: https://huggingface.co/datasets/cobrayyxx/FLEURS_ID-EN.audiotranslation1K<n<10K1 likes109 downloads2y agoHugging Face30ymoslem /FLEURS-GA-EN Dataset Details This is the Irish-to-English portion of the FLEURS dataset. Fleurs is the speech version of the FLoRes machine translation benchmark. The Irish portion consists of 3991 utterances, which correspond to approximately 16 hours and 45 minutes (16:45:17) of audio data. Dataset Structure DatasetDict({ train: Dataset({ features: ['id', 'audio', 'text_ga', 'text_en'], num_rows: 3991 }) }) Citation @article{fleurs2022arxiv… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/FLEURS-GA-EN.audioautomatic-speech-recognition1K<n<10K1 likes107 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.