datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.fleursbelebele-fleurs
Belebele-Fleurs
Belebele-Fleurs is a dataset suitable to evaluate two core tasks:
Multilingual Spoken Language Understanding (Listening Comprehension): For each spoken paragraph, the task is to answer a multiple-choice question. The question and four answer choices are provided in text form.
Multilingual Long-Form Automatic Speech Recognition (ASR) with Diverse Speakers: By concatenating sentence-level utterances, long-form audio clips (ranging from 30 seconds to 1 minute 30… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/belebele-fleurs.cs-fleurs
🌍 CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
📖 Overview
CS-FLEURS is a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages.
113 unique code-switched language pairs across 52 languages
300 hours of speech data, both read and synthetic
📊 Dataset Statistics
CS-FLEURS consists of the following subsets:
Read-Test: 14 X-English pairs, read speech… See the full description on the dataset page: https://huggingface.co/datasets/byan/cs-fleurs.sib-fleurs
SIB-Fleurs
SIB-Fleurs is a dataset suitable to evaluate Multilingual Spoken Language Understanding. For each utterance in Fleurs, the task is to determine the topic the utterance belongs to.
The topics are:
Science/Technology
Travel
Politics
Sports
Health
Entertainment
Geography
Preliminary evaluations can be found at the bottom of the README. The preliminary results in full detail are available in ./results.csv*.
Dataset creation
This dataset processes and merges… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/sib-fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/JianShangQuan/fleurs.fleurs-full
FLEURS Full - Test Set for ASR Benchmarking
Complete test set of Google FLEURS for all 30 languages supported by Qwen3-ASR, prepared for benchmarking with FluidAudio.
Languages (30)
Asian Languages (13)
Code
Language
Samples
cmn_hans_cn
Chinese (Mandarin)
945
yue_hant_hk
Cantonese
819
ja_jp
Japanese
650
ko_kr
Korean
382
vi_vn
Vietnamese
857
th_th
Thai
1,021
id_id
Indonesian
687
ms_my
Malay
749
hi_in
Hindi
418
ar_eg
Arabic (Egyptian)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs-full.sib-fleurs-multilingual-minifleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/czqdfsdf/fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ItzmeNishh/fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/yohannabelay/fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ssschuang/fleurs.indo-cv17-titml-fleursfleurs-farsi
FLEURS Farsi (fa_ir) - Processed Dataset
Dataset Description
This dataset contains the Farsi (Persian, fa_ir) portion of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) dataset, processed into a Hugging Face datasets compatible format. FLEURS is a many-language speech dataset created by Google, designed for evaluating speech recognition systems, particularly in low-resource scenarios.
This version includes audio recordings and their… See the full description on the dataset page: https://huggingface.co/datasets/MohammadGholizadeh/fleurs-farsi.fleurs-r-neucodec-all-languages
FLEURS-R NeuCodec All Languages
FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102
locales, plus a speaker label FLEURS itself does not ship.
Layout
data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the
viewer shows).
audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members
named audio/{locale}/{split}/{id}.wav (the path column).
neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.fleurs
FLEURS Test Dataset
Reorganized FLEURS test dataset with audio and transcripts together.
Structure
fleurs-test/
├── en_us/
│ ├── en_us_0000.wav
│ ├── en_us_0001.wav
│ ├── ...
│ ├── en_us.trans.txt (LibriSpeech format)
│ ├── en_us.csv (detailed metadata)
│ └── en_us.json (JSON metadata)
├── fr_fr/
│ └── ...
└── ...
Languages
bg_bg: 350 test samples
cs_cz: 350 test samples
da_dk: 930 test samples
de_de: 350 test samples
el_gr: 650… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs.cntt2-fleurs
Dataset Card for "cntt2-fleurs"
More Information needed
mixed-language-detection-pilot-fleurs-voices
Mixed-Language Speech Detection Pilot — Native FLEURS Voices
This is the native-reference revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed in this revision
Synthetic speech is cloned from 36 real… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-fleurs-voices.fleurs_test
FLEURS Test Dataset with Enhanced Metadata
This dataset is an enhanced version of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) test set, restructured with complete metadata for easier use in automatic speech recognition (ASR) and multilingual speech processing tasks.
Dataset Description
FLEURS is a multilingual speech benchmark dataset designed to evaluate universal speech representations. This particular version focuses on 25 European… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/fleurs_test.fleurs-ethiopian-v2
FLEURS — Ethiopian Languages
This dataset is a restructured v2 conversion of the Google FLEURS dataset for two Ethiopian languages: Amharic (am_et) and Oromo (om_et).
Subsets
Subset
Language
ISO 639-2
Train
Dev
Test
amh
Amharic
amh
3,163
223
516
orm
Oromo
orm
1,701
19
41
Splits
Split
Description
train
Training split
dev
Development split (renamed from validation in original FLEURS)
test
Test split
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/fleurs-ethiopian-v2.fleursfleurs-vi-preprocessed-v2fleurs-ro
FLEURS-RO (Rich Orthography)
Test-only Indic rich-transcription benchmark derived from google/fleurs. Each reference transcript is regenerated with grammatical punctuation, formatted numerals, and Indic-script orthographic conventions through an LLM curation pipeline whose prompts were iteratively refined against native-speaker review.
Released alongside the SCRIBE evaluation framework in SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR (accepted at… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/fleurs-ro.Fleurs_Irish_normalizedfleurs-textgridsThis dataset provides TextGrids with tiers phones in IPA and words in usual script corresponding to field word_segmented in
mahesh27/fleurs-ipa.
Alignments are generated using mahesh27/mms-300m-ipa-fleurs along with post silence trimming as per the paper.
Usage
Download textgrids.zip and extract such that the directory structure looks like textgrids/en_us/123456789.TextGrid. For metadata, load as usual:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/mahesh27/fleurs-textgrids.ru_common_voice_sova_rudevices_golos_fleursfleurs-badini
FLEURS-Badini
Dataset Summary
FLEURS-Badini is a speech dataset for the Badini dialect of Northern Kurdish, designed for research in:
Automatic Speech Recognition (ASR)
Speech-to-Text Translation (S2TT)
It is a dialect-specific extension of the FLEURS benchmark, providing aligned speech–text–translation data for a low-resource language variant.
The dataset contains 5,224 utterances (~15h40m) recorded from 45 speakers.
Supported Tasks
Automatic Speech… See the full description on the dataset page: https://huggingface.co/datasets/BadiniSpeechNLP/fleurs-badini.fleurs_mkFLEURS_ID-EN
Dataset Details
This is the Indonesia-to-English dataset for Speech Translation task. This dataset is acquired from FLEURS.
Fleurs is the speech version of the FLoRes machine translation benchmark. Fleurs has many languages, one of which is Indonesia for about 3561 utterances and approximately 12 hours and 24 minutes of audio data.
Processing Steps
Before the Fleurs dataset is extracted, there are some preprocessing steps to the data:
Remove some unused columns (since we… See the full description on the dataset page: https://huggingface.co/datasets/cobrayyxx/FLEURS_ID-EN.FLEURS-GA-EN
Dataset Details
This is the Irish-to-English portion of the FLEURS dataset.
Fleurs is the speech version of the FLoRes machine translation benchmark.
The Irish portion consists of 3991 utterances, which correspond to approximately 16 hours and 45 minutes (16:45:17) of audio data.
Dataset Structure
DatasetDict({
train: Dataset({
features: ['id', 'audio', 'text_ga', 'text_en'],
num_rows: 3991
})
})
Citation
@article{fleurs2022arxiv… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/FLEURS-GA-EN.
