datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
French_Documents_Dataset_PDF
French Documents Dataset (PDF)
This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.french-tts-conversational-dataset
French Conversational TTS Dataset
Dataset Description
This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's Voxtral Mini TTS model (voxtral-mini-tts-2603). It covers three B2B industry verticals with balanced male/female speaker distribution.
Verticals
Vertical
Description
fintech_banking
Banking operations, account inquiries, fraud alerts, investments, customer service
ecommerce_logistics
Order… See the full description on the dataset page: https://huggingface.co/datasets/JDKdev/french-tts-conversational-dataset.Emilia-dataset-french-with-gender
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.Emilia-dataset-french-splitcml_tts_dataset_frenchsakthivinash-Language_DetectionCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données https://huggingface.co/datasets/sakthivinash/Language_Detection.
french_tv_media_dataset_2026
Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus
Résumé (Abstract)
Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/Archime/french_tv_media_dataset_2026.french_tv_media_dataset_2026
Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus
Résumé (Abstract)
Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/madoss/french_tv_media_dataset_2026.croissant_french_datasetFrench-Alpaca-dataset-Instruct-110K110368 French instructions generated by OpenAI GPT-3.5-turbo in Alpaca Format to finetune general models
Created by Jonathan Pacifico, 2024Please credit my name if you use this dataset in your project.
english-french-datasetFrench-Alpaca-dataset-Instruct-55K55184 french instructions generated by OpenAI GPT-3.5
in Alpaca Format to finetune general models
Created by Jonathan Pacifico
license: apache-2.0
Please credit my name if you use this dataset in your project.
Wolof-to-French_Translation-Dataset
Dataset Wolof ↔ Français
🧩 Présentation
Ce dataset contient plus de 30 000 paires phrase Wolof – phrase Française.Chaque ligne est structurée comme suit :
Wolof (input)
Français (target)
Phrase en Wolof
Phrase correspondante en Français
Il a été conçu pour la traduction automatique et les tâches de NLP impliquant le Wolof et le Français.
📚 Provenance et nettoyage
Le dataset a été créé en compilant différentes sources accessibles… See the full description on the dataset page: https://huggingface.co/datasets/MaroneAI/Wolof-to-French_Translation-Dataset.Makxxx-french_CEFRCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données huggingface.co/datasets/Makxxx/french_CEFR.
vekkt-french_CEFRCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données huggingface.co/datasets/vekkt/french_CEFR.
french-tts-conversational-dataset
French Conversational TTS Dataset
Dataset Description
This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's Voxtral Mini TTS model (voxtral-mini-tts-2603). It covers three B2B industry verticals with balanced male/female speaker distribution.
Verticals
Vertical
Description
fintech_banking
Banking operations, account inquiries, fraud alerts, investments, customer service
ecommerce_logistics
Order… See the full description on the dataset page: https://huggingface.co/datasets/voxozi/french-tts-conversational-dataset.quebecois_canadian_french_datasetfrench-instruction-datasetenglish-wolof-french-datasetfrench-speech-datasetjfv-french-style-conditioning-dataset-v1.0
JFV French Paired Style-Conditioning Dataset
At a Glance
Item
Value
Language
French
Source
Single-author blog corpus, 2005–2025
Public release
v1.0
Aligned units in public release
1,484
Texts in public aligned release
7,420
Original experiment
1,492 aligned units / 7,460 texts
Generated conditions
Ministral baseline; profile; profile + five-shot examples
Primary use
Paired study of stylistic conditioning and evaluation-metric validity… See the full description on the dataset page: https://huggingface.co/datasets/PeggyVallin/jfv-french-style-conditioning-dataset-v1.0.french-call-center-speech-dataset
French Call Center Speech Dataset: 1,000+ Hours with Transcripts
1,000+ hours of real-world French call center audio with transcripts. Train speech recognition, sentiment analysis, and customer support AI models on authentic telephone conversations
Dataset Summary
Key Features
✅ 1,000+ hours of inbound & outbound calls✅ 100% French telephone conversations✅ Real-world audio - no synthetic data✅ Full transcripts in French and in English
Full… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/french-call-center-speech-dataset.facebook-community-alignment-dataset_french_dpo
Description
This is the Community Alignment dataset which we've cleaned up to keep only the French datas (+ deduplication) and reformatted for DPO finetuning.For more details on the dataset itself, please consult the original dataset card or the paper.
Original authors
@article{zhang2025cultivating,
title = {Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset},
author = {Lily Hong Zhang and Smitha Milli and Karen Jusko and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/facebook-community-alignment-dataset_french_dpo.French_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 31,106 hours of processed French (FR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/French_Call_Center_Audio_Dataset_Dual_Channel.facebook-community-alignment-dataset_french_conversation
Description
This is the Community Alignment dataset which we've cleaned up to keep only the French datas (+ deduplication) and reformatted as a conversation to simplify his use for alignment finetuning.For more details on the dataset itself, please consult the original dataset card or the paper.
Original authors
@article{zhang2025cultivating,
title = {Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset},
author = {Lily Hong Zhang… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/facebook-community-alignment-dataset_french_conversation.query-intent-detection-dataset-frenchmistral-legal-french-dataset
Mistral Legal French Dataset
A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy.
📋 Table of Contents
Overview
Dataset Composition
Methodology
1. Chain-of-Thought Generation
2. LegalKit Extraction
3. Curriculum Learning Fusion
Data Format
Quality Metrics
Usage
Citations
License
🎯 Overview
This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two… See the full description on the dataset page: https://huggingface.co/datasets/davidpistori/mistral-legal-french-dataset.French-Speech-Dataset
🎧 French Speech Dataset
The French Speech Dataset is a comprehensive speech audio dataset designed to deliver high-quality and diverse audio data for advanced AI and machine learning applications. It includes 198 hours of audio data across 912 files, provided in MP3 and WAV formats, with a total size of 445 MB. This well-structured audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a wide age distribution from 18 to 50+ years.… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/French-Speech-Dataset.DILA_FRENCH_DATASETsmol-smoltalk-french-instruction-dataset
French Instruction Dataset for Nanochat
This dataset is a transformed version of vonewman/french-instruction-dataset.
It has been formatted to match the SmolTalk structure required by nanochat.
Format
Format: ChatML / SmolTalk
Column: messages (List of dicts with role and content)
Usage with Nanochat
self.ds = load_dataset("Volko76/smol-smoltalk-french-instruction-dataset", split=split)
