datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
French-PD-Books
🇫🇷 French Public Domain Books 🇫🇷
French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.French-PD-Newspapers
🇫🇷 French Public Domain Newspapers 🇫🇷
French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.french_CEFRTransWeb-Edu-FrenchFrench-Science-Commons
French Science Commons
French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres.
Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.French_Documents_Dataset_PDF
French Documents Dataset (PDF)
This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.hatvp_raw_archivesfrench_CEFRfrench-30b
Dataset Card for "french_30b2"
More Information needed
SPEEED_s3_words_french_100k-300kfrench-tts-conversational-dataset
French Conversational TTS Dataset
Dataset Description
This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's Voxtral Mini TTS model (voxtral-mini-tts-2603). It covers three B2B industry verticals with balanced male/female speaker distribution.
Verticals
Vertical
Description
fintech_banking
Banking operations, account inquiries, fraud alerts, investments, customer service
ecommerce_logistics
Order… See the full description on the dataset page: https://huggingface.co/datasets/JDKdev/french-tts-conversational-dataset.french-conversation+15 hours of speech data from TTS and text file recording.
+9k utterances from various sources, novels, parliamentary debates, professional language.
mt-bench-french
MT-Bench-French
This is a French version of MT-Bench, created to evaluate the multi-turn conversation and instruction-following capabilities of LLMs. Similar to its original version, MT-Bench-French comprises 80 high-quality, multi-turn questions spanning eight main categories.
All questions have undergone translation into French and thorough human review to guarantee the use of suitable and authentic wording, meaningful content for assessing LLMs' capabilities in the French… See the full description on the dataset page: https://huggingface.co/datasets/bofenghuang/mt-bench-french.French-PD-diverse43,085,129,931 words
french-dialogue-tts-1000hFRENCH-ONLY-Common-Crawl-2026-25SPEEED_s3_words_french_0k-100k
Multilingual Audio Alignments - Processed (Mixed Text/Phonemes)
This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (french).
Curriculum Learning
This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule:
p_start: 0.0 (starting probability of using phonemes)
p_end: 0.0 (ending probability of using phonemes)
curriculum_rows: 400000 (rows over which probability increases)
Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_french_0k-100k.CXM_Arena_French
Dataset Card for CXM Arena French Benchmark Suite
Dataset Description
This dataset, "CXM Arena French Benchmark Suite," is a comprehensive collection designed to evaluate various AI capabilities within the Customer Experience Management (CXM) domain, specifically for the French language. It is closely modeled after the original CXM_Arena benchmark, but all data is in French. The suite consolidates five distinct tasks into a unified benchmark, enabling robust testing of… See the full description on the dataset page: https://huggingface.co/datasets/sprinklr-huggingface/CXM_Arena_French.french_instruct
🧑🏫 French Instruct
The French Instruct dataset is a collection of instructions with their corresponding answers (sometimes multi-turn conversations) entirely in French. The dataset is also available on GitHub.
📊 Overview
The dataset is composed of 276K conversations between a user and an assistant for a total of approximately 85M tokens.
I also added annotations for each document to indicate if it was generated or written by a human, the style of… See the full description on the dataset page: https://huggingface.co/datasets/angeluriot/french_instruct.caption-wit_base_french
Description
This dataset is a processed version of wikimedia/wit_base.We converted the images to PIL format and kept only the lines containing French.The French texts come either from the caption_attribution_description column of the original dataset which we have renamed caption here, or from the wit_features column which we have renamed descriptions.The descriptions column is a list because it can contain several texts itself.
For example, line 8 of the dataset contains three… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/caption-wit_base_french.French-PD-diverse43,085,129,931 words
Emilia-dataset-french-splitMMLU_FrenchFrench version of MMLU dataset tranlasted by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT.
french_5p
Dataset Card for "french_5p"
More Information needed
Emilia-dataset-french-with-gender
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.SLR-Bench-French
🧠 SLR-Bench-French: Scalable Logical Reasoning Benchmark (French Edition)
SLR-Bench Multilingual Versions:
SLR-Bench-French is the French-language pendant of the original SLR-Benchdataset.
It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into French.
This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-French.ddxplus-french
Dataset Description
We are releasing under the CC-BY licence a new large-scale dataset for Automatic Symptom Detection (ASD) and Automatic Diagnosis (AD) systems in the medical domain. The dataset contains patients synthesized using a proprietary medical knowledge base and a commercial rule-based AD system. Patients in the dataset are characterized by their socio-demographic data, a pathology they are suffering from, a set of symptoms and antecedents related to this pathology, and a… See the full description on the dataset page: https://huggingface.co/datasets/aai530-group6/ddxplus-french.cml_tts_dataset_frencheighteenth_century_french_novels
General information
This dataset contains 12 Mio Token of Literary French prose 1751-1800 in plain text format, built within the project 'Mining and Modeling Text' (2019-2023) at Trier University.
For the dataset in XML/TEI see the GitHub repository of the project.
Collection de romans français du dix-huitième siècle (1751-1800) / Collection of Eighteenth-Century French Novels (1751-1800)
This collection of Eighteenth-Century French Novels contains 200 digital… See the full description on the dataset page: https://huggingface.co/datasets/roettger/eighteenth_century_french_novels.french-dialogue-tts-100h
