datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.fleursbelebele-fleurs
Belebele-Fleurs
Belebele-Fleurs is a dataset suitable to evaluate two core tasks:
Multilingual Spoken Language Understanding (Listening Comprehension): For each spoken paragraph, the task is to answer a multiple-choice question. The question and four answer choices are provided in text form.
Multilingual Long-Form Automatic Speech Recognition (ASR) with Diverse Speakers: By concatenating sentence-level utterances, long-form audio clips (ranging from 30 seconds to 1 minute 30… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/belebele-fleurs.cs-fleurs
🌍 CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
📖 Overview
CS-FLEURS is a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages.
113 unique code-switched language pairs across 52 languages
300 hours of speech data, both read and synthetic
📊 Dataset Statistics
CS-FLEURS consists of the following subsets:
Read-Test: 14 X-English pairs, read speech… See the full description on the dataset page: https://huggingface.co/datasets/byan/cs-fleurs.sib-fleurs
SIB-Fleurs
SIB-Fleurs is a dataset suitable to evaluate Multilingual Spoken Language Understanding. For each utterance in Fleurs, the task is to determine the topic the utterance belongs to.
The topics are:
Science/Technology
Travel
Politics
Sports
Health
Entertainment
Geography
Preliminary evaluations can be found at the bottom of the README. The preliminary results in full detail are available in ./results.csv*.
Dataset creation
This dataset processes and merges… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/sib-fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/JianShangQuan/fleurs.sib-fleurs-multilingual-minifleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/czqdfsdf/fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ssschuang/fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ItzmeNishh/fleurs.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/yohannabelay/fleurs.indo-cv17-titml-fleurspreprocessed_jsut_jsss_css10_fleurs_common_voice_11
Dataset Card for "preprocessed_jsut_jsss_css10_fleurs_common_voice_11"
More Information needed
fleurs-farsi
FLEURS Farsi (fa_ir) - Processed Dataset
Dataset Description
This dataset contains the Farsi (Persian, fa_ir) portion of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) dataset, processed into a Hugging Face datasets compatible format. FLEURS is a many-language speech dataset created by Google, designed for evaluating speech recognition systems, particularly in low-resource scenarios.
This version includes audio recordings and their… See the full description on the dataset page: https://huggingface.co/datasets/MohammadGholizadeh/fleurs-farsi.fleurs-r-neucodec-all-languages
FLEURS-R NeuCodec All Languages
FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102
locales, plus a speaker label FLEURS itself does not ship.
Layout
data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the
viewer shows).
audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members
named audio/{locale}/{split}/{id}.wav (the path column).
neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.cntt2-fleurs
Dataset Card for "cntt2-fleurs"
More Information needed
fleurs-neucodec
Dataset
Dataset Statistics
This table shows the number of examples per language configuration and split.
config_name
train_examples
validation_examples
test_examples
af_za
1.032
198
264
am_et
3.163
223
516
ar_eg
2.104
295
428
as_in
2.812
418
984
ast_es
2.511
398
946
az_az
2.665
400923
be_by
2.433
408
967
bg_bg
2.973
395
658
bn_in
3.006
402
920
bs_ba
3.091
400
925
ca_es
2.300
404
940
ceb_ph
3.261
225
541
ckb_iq
3.040
386
922
cmn_hans_cn… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/fleurs-neucodec.mixed-language-detection-pilot-fleurs-voices
Mixed-Language Speech Detection Pilot — Native FLEURS Voices
This is the native-reference revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed in this revision
Synthetic speech is cloned from 36 real… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-fleurs-voices.fleurs-ethiopian-v2
FLEURS — Ethiopian Languages
This dataset is a restructured v2 conversion of the Google FLEURS dataset for two Ethiopian languages: Amharic (am_et) and Oromo (om_et).
Subsets
Subset
Language
ISO 639-2
Train
Dev
Test
amh
Amharic
amh
3,163
223
516
orm
Oromo
orm
1,701
19
41
Splits
Split
Description
train
Training split
dev
Development split (renamed from validation in original FLEURS)
test
Test split
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/fleurs-ethiopian-v2.fleurs_test
FLEURS Test Dataset with Enhanced Metadata
This dataset is an enhanced version of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) test set, restructured with complete metadata for easier use in automatic speech recognition (ASR) and multilingual speech processing tasks.
Dataset Description
FLEURS is a multilingual speech benchmark dataset designed to evaluate universal speech representations. This particular version focuses on 25 European… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/fleurs_test.fleurs-ipaThe dataset is intended to be used along side google/fleurs where both id and audio_file have to be matched (FLEURS may contain path from which file name has to be matched).
Citation
@article{akavarapu2026phoneme,
title={Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation},
author={Akavarapu, V.S.D.S.Mahesh and Daniel, Michael and J{\"a}ger, Gerhard},
year={2026},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/mahesh27/fleurs-ipa.fleursfleurs-vi-preprocessed-v2Fleurs_Irish_normalizedfleurs-ro
FLEURS-RO (Rich Orthography)
Test-only Indic rich-transcription benchmark derived from google/fleurs. Each reference transcript is regenerated with grammatical punctuation, formatted numerals, and Indic-script orthographic conventions through an LLM curation pipeline whose prompts were iteratively refined against native-speaker review.
Released alongside the SCRIBE evaluation framework in SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR (accepted at… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/fleurs-ro.fleurs-textgridsThis dataset provides TextGrids with tiers phones in IPA and words in usual script corresponding to field word_segmented in
mahesh27/fleurs-ipa.
Alignments are generated using mahesh27/mms-300m-ipa-fleurs along with post silence trimming as per the paper.
Usage
Download textgrids.zip and extract such that the directory structure looks like textgrids/en_us/123456789.TextGrid. For metadata, load as usual:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/mahesh27/fleurs-textgrids.ru_common_voice_sova_rudevices_golos_fleursfleurs-badini
FLEURS-Badini
Dataset Summary
FLEURS-Badini is a speech dataset for the Badini dialect of Northern Kurdish, designed for research in:
Automatic Speech Recognition (ASR)
Speech-to-Text Translation (S2TT)
It is a dialect-specific extension of the FLEURS benchmark, providing aligned speech–text–translation data for a low-resource language variant.
The dataset contains 5,224 utterances (~15h40m) recorded from 45 speakers.
Supported Tasks
Automatic Speech… See the full description on the dataset page: https://huggingface.co/datasets/BadiniSpeechNLP/fleurs-badini.ovos-stt-bench-fleurs-ca-ES
OVOS stt bench — fleurs-ca-ES
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
google/fleurs.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-fleurs-ca-ES.FLEURS-GA-EN
Dataset Details
This is the Irish-to-English portion of the FLEURS dataset.
Fleurs is the speech version of the FLoRes machine translation benchmark.
The Irish portion consists of 3991 utterances, which correspond to approximately 16 hours and 45 minutes (16:45:17) of audio data.
Dataset Structure
DatasetDict({
train: Dataset({
features: ['id', 'audio', 'text_ga', 'text_en'],
num_rows: 3991
})
})
Citation
@article{fleurs2022arxiv… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/FLEURS-GA-EN.
