datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-dataset-french-with-gender
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.common_voice_13_french_phoneme
Common Voice 13 French Phoneme
Dataset Summary
This dataset is a curated version of the French subset of Common Voice 13.0, enriched with a phonetic transcription column (phoneme).
It was created by the Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) to support research in speech processing, specifically for tasks requiring phonetic alignment, phoneme recognition, and robust speech-to-text applications in French.
The dataset retains the… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/common_voice_13_french_phoneme.french_tv_media_dataset_2026
Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus
Résumé (Abstract)
Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/Archime/french_tv_media_dataset_2026.SIWIS_French_Speech_Synthesis_Database
SIWIS French Speech Synthesis Database
This README provides a concise description of the dataset, including its structure, file naming conventions, and known labeling issues. Additionally, suggestions for potential improvements are outlined in the TODO section.
The dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting its use for any purpose.
For more details about the database design and recording process, please refer… See the full description on the dataset page: https://huggingface.co/datasets/Aviv-anthonnyolime/SIWIS_French_Speech_Synthesis_Database.french_tv_media_dataset_2026
Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus
Résumé (Abstract)
Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/madoss/french_tv_media_dataset_2026.global-french-speech
Global French Speech Dataset
Speaker-attributed spontaneous speech. 50 labelled contributors across 11 self-reported origin varieties, recorded on their own devices in their own environments. Every clip carries origin region / variety, mother tongue, gender, device, OS, recording environment.
Hours
0.54
Clips
50
Speakers
50
Origin varieties
11
Languages
1
Configs
2
Speaker metadata
origin region / variety, mother tongue, gender, device, OS, recording… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/global-french-speech.wolof-french-asr
Wolof-French ASR Dataset
Description
Dataset unifié pour l'entraînement de modèles de reconnaissance automatique de la parole (ASR) en wolof et français. Le wolof est une langue d'Afrique de l'Ouest parlée principalement au Sénégal par plus de 10 millions de locuteurs. Les locuteurs wolof pratiquent fréquemment le code-switching (alternance wolof/français), ce qui rend indispensable un modèle ASR capable de transcrire les deux langues.
Composition du dataset… See the full description on the dataset page: https://huggingface.co/datasets/serge-wilson/wolof-french-asr.french-education-speech
French Education Speech - Transcribed Dataset
High-quality French educational speech dataset transcribed with OpenAI Whisper API, prepared for training automatic speech recognition (ASR) models.
Dataset Summary
This dataset contains 3,933 transcribed audio segments from the French educational domain, totaling approximately 12.82 hours of audio. All transcriptions were performed using OpenAI Whisper API (optimized Whisper-1 model) to ensure maximum accuracy, especially for… See the full description on the dataset page: https://huggingface.co/datasets/MEscriva/french-education-speech.multilingual_librispeech_french_phoneme
Multilingual LibriSpeech French Phoneme
Dataset Summary
This dataset is a curated version of the French subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into French acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_french_phoneme.multilingual_librispeech_french_punctuated
Multilingual LibriSpeech French, punctuated and capitalized (train)
A derivative of the French part of Multilingual LibriSpeech (MLS), the corpus of read audiobooks from LibriVox published by Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve and Ronan Collobert (Facebook AI Research). MLS distributes its transcriptions lowercased and without any punctuation. This dataset keeps that upstream text untouched in the text column and adds a second column, text_punctuated… See the full description on the dataset page: https://huggingface.co/datasets/Ugiat/multilingual_librispeech_french_punctuated.french-education-speech
French Education Speech - Transcribed Dataset
High-quality French educational speech dataset transcribed with OpenAI Whisper API, prepared for training automatic speech recognition (ASR) models.
Dataset Summary
This dataset contains 3,933 transcribed audio segments from the French educational domain, totaling approximately 12.82 hours of audio. All transcriptions were performed using OpenAI Whisper API (optimized Whisper-1 model) to ensure maximum accuracy, especially for… See the full description on the dataset page: https://huggingface.co/datasets/Lexia-Labs/french-education-speech.French_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 31,106 hours of processed French (FR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/French_Call_Center_Audio_Dataset_Dual_Channel.French-Speech-Dataset
🎧 French Speech Dataset
The French Speech Dataset is a comprehensive speech audio dataset designed to deliver high-quality and diverse audio data for advanced AI and machine learning applications. It includes 198 hours of audio data across 912 files, provided in MP3 and WAV formats, with a total size of 445 MB. This well-structured audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a wide age distribution from 18 to 50+ years.… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/French-Speech-Dataset.french-long-form-testglobal-french-speech_structured
Zeldeo/global-french-speech_structured
Dataset ASR restructuré depuis SilencioNetwork/global-french-speech
(config=french_canada, split=train).
Nombre d'exemples : 25.
Métadonnées ajoutées : source_dataset, type, langue_accent.
Normalisation texte : aucune.
Colonnes conservées
audio
gender
dialect
emotions
language
location
noise_sources
transcript
age_band
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Zeldeo/global-french-speech_structured.french-conversation_structured
Zeldeo/french-conversation_structured
Dataset ASR restructuré depuis Snit/french-conversation
(config=default, split=train).
Nombre d'exemples : 98.
Métadonnées ajoutées : source_dataset, type, langue_accent.
Normalisation texte : aucune.
Colonnes conservées
audio
transcription
id
part
audio_path
Usage
from datasets import load_dataset
ds = load_dataset("Zeldeo/french-conversation_structured", split="train")
print(ds[0])
French-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 31,106 hours of processed French (FR) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/French-Call-Center-Audio-Dataset-Single-Channel.french-call-center-speech-fr
French Call Center Speech Dataset (~164 hours)
DescriptionA large-scale dataset of real French call center recordings.Telephone-quality conversations between customers and agents.Useful for ASR, NLP, speech analytics, and training voicebots.
Technical details
Language: French
Total duration: ~164 hours
Format: MP3
Channels: Mono (1 channel, mixed client and agent)
Sample rate: 8000 Hz
Bitrate: 32 kbps
Metadata: Not available
LicenseCommercial license only.… See the full description on the dataset page: https://huggingface.co/datasets/MaratDV/french-call-center-speech-fr.french_homophone_asrThe dataset, created and used in Mohebbi et al., (2023), includes instances of homophones in French spoken language, where an ASR model has to attend to syntactic cues in the context to disambiguate spoken words with identical pronunciations for transcription.
The test set (fr) of the Common Voice 11.0 (Ardila et al., 2020) is used to discover instances of three specific grammatical syntactic templates in which homophony may appear.
For the purpose of analysis, examples are filtered so that… See the full description on the dataset page: https://huggingface.co/datasets/hosein-m/french_homophone_asr.french-speech-recognition-dataset
French Speech Dataset for recognition task
Dataset comprises 547 hours of telephone dialogues in French, collected from 964 native speakers across various topics and domains, with an impressive 98% Word Accuracy Rate. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/french-speech-recognition-dataset.french-speech-samples
French Speech Samples
This sample shows French contributor speech paired with source transcripts. It is meant to help buyers review recording quality, transcript alignment, and metadata structure before scoping a larger delivery.
What This Shows
French single-speaker recordings
Transcript alignment from the source dataset
Clip-level metadata for format and review context
Dataset Specifications
Field
Value
Modality
Audio
Language… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/french-speech-samples.YodaLingua-French
YodaLingua-French
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the French portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
56,965 audio–transcription pairs
Total duration
171 hours
Speakers
2,320 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-French.
