datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.fleursfleurs-rbelebele-fleurs
Belebele-Fleurs
Belebele-Fleurs is a dataset suitable to evaluate two core tasks:
Multilingual Spoken Language Understanding (Listening Comprehension): For each spoken paragraph, the task is to answer a multiple-choice question. The question and four answer choices are provided in text form.
Multilingual Long-Form Automatic Speech Recognition (ASR) with Diverse Speakers: By concatenating sentence-level utterances, long-form audio clips (ranging from 30 seconds to 1 minute 30… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/belebele-fleurs.cs-fleurs
🌍 CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
📖 Overview
CS-FLEURS is a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages.
113 unique code-switched language pairs across 52 languages
300 hours of speech data, both read and synthetic
📊 Dataset Statistics
CS-FLEURS consists of the following subsets:
Read-Test: 14 X-English pairs, read speech… See the full description on the dataset page: https://huggingface.co/datasets/byan/cs-fleurs.sib-fleurs
SIB-Fleurs
SIB-Fleurs is a dataset suitable to evaluate Multilingual Spoken Language Understanding. For each utterance in Fleurs, the task is to determine the topic the utterance belongs to.
The topics are:
Science/Technology
Travel
Politics
Sports
Health
Entertainment
Geography
Preliminary evaluations can be found at the bottom of the README. The preliminary results in full detail are available in ./results.csv*.
Dataset creation
This dataset processes and merges… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/sib-fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/JianShangQuan/fleurs.sib-fleurs-multilingual-minifleurs-full
FLEURS Full - Test Set for ASR Benchmarking
Complete test set of Google FLEURS for all 30 languages supported by Qwen3-ASR, prepared for benchmarking with FluidAudio.
Languages (30)
Asian Languages (13)
Code
Language
Samples
cmn_hans_cn
Chinese (Mandarin)
945
yue_hant_hk
Cantonese
819
ja_jp
Japanese
650
ko_kr
Korean
382
vi_vn
Vietnamese
857
th_th
Thai
1,021
id_id
Indonesian
687
ms_my
Malay
749
hi_in
Hindi
418
ar_eg
Arabic (Egyptian)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs-full.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/czqdfsdf/fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ssschuang/fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ItzmeNishh/fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/yohannabelay/fleurs.fleurs_tts_dump_24000indo-cv17-titml-fleurswav_to_vec_common_voice_fleurs_without_diacs
Dataset Card for "wav_to_vec_common_voice_fleurs_without_diacs"
More Information needed
fleurs-hs
FLEURS-HS
An extension of the FLEURS dataset for synthetic speech detection using text-to-speech, featured in the paper Synthetic speech detection with Wav2Vec 2.0 in various language settings.
This dataset is 1 of 3 used in the paper, the others being:
FLEURS-HS VITS
test set containing (generally) more difficult synthetic samples
separated due to different licensing
ARCTIC-HS
extension of the CMU_ARCTIC and L2-ARCTIC sets in a similar manner
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/realnetworks-kontxt/fleurs-hs.preprocessed_jsut_jsss_css10_fleurs_common_voice_11
Dataset Card for "preprocessed_jsut_jsss_css10_fleurs_common_voice_11"
More Information needed
fleurs-r-neucodec-all-languages
FLEURS-R NeuCodec All Languages
FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102
locales, plus a speaker label FLEURS itself does not ship.
Layout
data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the
viewer shows).
audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members
named audio/{locale}/{split}/{id}.wav (the path column).
neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/cahya/fleurs.fleurs_ipafleurs-farsi
FLEURS Farsi (fa_ir) - Processed Dataset
Dataset Description
This dataset contains the Farsi (Persian, fa_ir) portion of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) dataset, processed into a Hugging Face datasets compatible format. FLEURS is a many-language speech dataset created by Google, designed for evaluating speech recognition systems, particularly in low-resource scenarios.
This version includes audio recordings and their… See the full description on the dataset page: https://huggingface.co/datasets/MohammadGholizadeh/fleurs-farsi.google_fleurs_plus_common_voice_11_arabic_language
Dataset Card for "google_fleurs_plus_common_voice_11_arabic_language"
More Information needed
fleurs-regmix-webdataset
FLEURS RegMix WebDataset
Public, training-oriented WebDataset conversion of google/fleurs, pinned to source revision 70bb2e84b976b7e960aa89f1c648e09c59f894dd.
Layout
Each language is a RegMix cluster at data/<language>/*.tar. Shard names keep the source split, data/<language>/<language>-<split>-<index>.tar, so a training
mixture can be assembled without pulling the FLEURS evaluation splits into it. Every sample is a pair with the same key:
<key>.opus: mono Ogg… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/fleurs-regmix-webdataset.asr-khm-ddd-fleurs-openslr2cntt2-fleurs
Dataset Card for "cntt2-fleurs"
More Information needed
fleurs
FLEURS Test Dataset
Reorganized FLEURS test dataset with audio and transcripts together.
Structure
fleurs-test/
├── en_us/
│ ├── en_us_0000.wav
│ ├── en_us_0001.wav
│ ├── ...
│ ├── en_us.trans.txt (LibriSpeech format)
│ ├── en_us.csv (detailed metadata)
│ └── en_us.json (JSON metadata)
├── fr_fr/
│ └── ...
└── ...
Languages
bg_bg: 350 test samples
cs_cz: 350 test samples
da_dk: 930 test samples
de_de: 350 test samples
el_gr: 650… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs.fleurs-neucodec
Dataset
Dataset Statistics
This table shows the number of examples per language configuration and split.
config_name
train_examples
validation_examples
test_examples
af_za
1.032
198
264
am_et
3.163
223
516
ar_eg
2.104
295
428
as_in
2.812
418
984
ast_es
2.511
398
946
az_az
2.665
400923
be_by
2.433
408
967
bg_bg
2.973
395
658
bn_in
3.006
402
920
bs_ba
3.091
400
925
ca_es
2.300
404
940
ceb_ph
3.261
225
541
ckb_iq
3.040
386
922
cmn_hans_cn… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/fleurs-neucodec.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/nanexxx/fleurs.fleurs_mk
