datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
French-PD-Newspapers
🇫🇷 French Public Domain Newspapers 🇫🇷
French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.French-PD-Books
🇫🇷 French Public Domain Books 🇫🇷
French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.French-Science-Commons
French Science Commons
French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres.
Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.French-PD-diverse43,085,129,931 words
FRENCH-ONLY-Common-Crawl-2026-25CXM_Arena_French
Dataset Card for CXM Arena French Benchmark Suite
Dataset Description
This dataset, "CXM Arena French Benchmark Suite," is a comprehensive collection designed to evaluate various AI capabilities within the Customer Experience Management (CXM) domain, specifically for the French language. It is closely modeled after the original CXM_Arena benchmark, but all data is in French. The suite consolidates five distinct tasks into a unified benchmark, enabling robust testing of… See the full description on the dataset page: https://huggingface.co/datasets/sprinklr-huggingface/CXM_Arena_French.French-PD-diverse43,085,129,931 words
SLR-Bench-French
🧠 SLR-Bench-French: Scalable Logical Reasoning Benchmark (French Edition)
SLR-Bench Multilingual Versions:
SLR-Bench-French is the French-language pendant of the original SLR-Benchdataset.
It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into French.
This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-French.cold-french-law
Collaborative Open Legal Data (COLD) - French Law
COLD French Law is a dataset containing over 800 000 french law articles, filtered and extracted from France's LEGI dataset and formatted as a single CSV file.
This dataset focuses on articles (codes, lois, décrets, arrêtés ...) identified as currently applicable french law.
A large portion of this dataset comes with machine-generated english translations, provided by Casetext, Part of Thomson Reuters using OpenAI's GPT-4.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-french-law.french_book_reviews
Dataset Card for French book reviews
I-Dataset Summary
The majority of review datasets are in English. There are datasets in other languages, but not many. Through this work, I would like to enrich the datasets in the French language(my mother tongue with Arabic).The data was retrieved from two French websites: Babelio and Critiques LibresLike Wikipedia, these two French sites are made possible by the contributions of volunteers who use the Internet to share their… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/french_book_reviews.paraphrasing-french
Attribution
MTEB-format derivative of ismailiismail/paraphrasing_french. Query = phrase; corpus = paraphrase.
alpaca-french-mixtral
License & Attribution
MTEB-format derivative of AIffl/Alpaca_french_mixtral (French Alpaca, Mixtral-translated). Query = instruction; corpus = answer. Deterministically subsampled to ~10k. Licensed under Apache-2.0 (same as source).
IMF-Reports
IMF Technical Assistance Reports — Recommendation Process Corpus
A page-grounded research corpus of 780 IMF technical-assistance report
records. It contains source PDFs, layout-aware Markdown, page-level text, extracted
visuals, metadata, observations, recommendations, and labeled links between observations
and recommendations.
Required acknowledgement
All research, publications, datasets, models, applications, or other work derived from
this corpus should… See the full description on the dataset page: https://huggingface.co/datasets/FrenchCastle/IMF-Reports.Argimi-Legal-French-Jurisprudence
The ArGiMi French Jurisprudence Dataset
This dataset contains a comprehensive collection of French case law, sourced from the official archives of French jurisprudence. It is divided into three distinct subdivisions: Constitutional ("constit"), Administrative ("cetat"), and Judiciary ("juri").
This dataset was created for the ArGiMi project, an open-source initiative dedicated to promoting open data and knowledge sharing. The project is a collaborative effort between Giskard… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Legal-French-Jurisprudence.French-Patent-1981-2026-Clean
🇫🇷 Brevets français 1981–2026 — Clean 🇫🇷
Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet).
Format : Parquet, prêt pour chargement streaming / distribué.
Source
Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI).
Génération… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/French-Patent-1981-2026-Clean.fama_french_datafrench-fiction-16-18th-century
French Fiction of the 16th–18th Centuries
A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model.
The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction.
Structure
Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.french-dictionary
French Dictionary
A ready-to-use offline French language dictionary derived from the French Wiktionary. Available in two formats to suit different use cases: SQLite for desktop applications and real-time querying, and Parquet for data science and machine learning pipelines.
Contains nearly 900,000 distinct word forms including conjugated verb forms, with structured definitions, usage examples, and rich linguistic metadata.
Acknowledgements
This dataset would not… See the full description on the dataset page: https://huggingface.co/datasets/Kartmaan/french-dictionary.french-administrative-hierarchy-data-quality
French Administrative Hierarchy Data Quality Benchmark
A reproducible benchmark for evaluating the validation, classification, and repair of French administrative geographic records.
The dataset is derived from the INSEE Code officiel géographique (COG) 2026 and focuses on the hierarchical relationship:
Region → Department → Commune
Important: commune_code is an INSEE/COG administrative identifier, not a postal code.This dataset is not intended for postal-address validation or… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/french-administrative-hierarchy-data-quality.isora-tax-administration
ISORA — International Survey on Revenue Administration, FY2014–FY2024
Every published answer of every ISORA survey round, in one clean long-format panel, with the
metadata you need to use it responsibly: what each question means in each questionnaire
generation, which questions changed wording (or meaning) between rounds, which jurisdictions
answered which question in which year, and how published values were revised between releases.
ISORA is the joint survey of national tax… See the full description on the dataset page: https://huggingface.co/datasets/FrenchCastle/isora-tax-administration.rte3-french
Dataset Card for French RTE-3
Dataset Summary
The RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge.
Like its English counterpart, the French RTE-3 dataset is composed of a development set and a test set, each containing 800 T/H pairs.
All T/H pairs were manually translated into French and proofread.
It is annotated for a 3-way task.
Please refer to this repository, if you want to use this French version… See the full description on the dataset page: https://huggingface.co/datasets/maximoss/rte3-french.hatecheck-french
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-french.English_French_Songs_Lyrics_Translation_Original
Original Songs Lyrics with French Translation
Dataset Summary
Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French.
Details of the number of songs by language of origin can be found in the table below:
Original language
Number of songs
en
75786
fr
18486
es
1743
it
803
de
691
sw
529
ko
193
id
169
pt
142
no
122
fi
113
sv
70
hr
53
so
43
ca
41
tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.clinical-qa-french-v2valoris-french-real-estate-prices
French Real Estate Prices — VALORIS Observatory
Aggregated real estate median prices (€/m²) computed from the French open data source DVF (Demandes de Valeurs Foncières) published by DGFiP.
Covers 93 departments and their communes of metropolitan France (excluding Alsace-Moselle departments 57, 67, 68 — local Livre Foncier system).
🔗 Interactive visualization & drill-down: valoris-immo.fr/observatoire
🏠 Publisher homepage: valoris-immo.fr
📊 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/VALORISIMMO/valoris-french-real-estate-prices.erudit-french-philosophy
Dataset Card for Dataset Name
Dataset Description
Dataset Summary
This dataset contains all french philosophy that has been published on erudit.org. It has been generated using a Bs4 web parser that you can find in this repo: https://github.com/MFGiguere/french-philosophy-generator.
Supported Tasks and Leaderboards
This dataset could be useful for this (non-exhaustive) set of tasks: detect if a text is philosophical or not, generate philosophical… See the full description on the dataset page: https://huggingface.co/datasets/mfgiguere/erudit-french-philosophy.banque-fr-2311
Dataset Card for "banque fr websites - 2311"
Dataset extracted from public websites by wordslab-webscraper in 2311:
domain: banque
language: fr
license: Apache 2.0
Dataset Sources
wordslab-webscraper follows the industry best practices for polite web scraping:
clearly identifies itself as a known text indexing bot: "bingbot"
doesn't try to hide the user IP address behind proxies
doesn't try to circumvent bots protection solutions
waits for a minimum delay between two… See the full description on the dataset page: https://huggingface.co/datasets/frenchtext/banque-fr-2311.SLR-Bench-French
🧠 SLR-Bench-French: Scalable Logical Reasoning Benchmark (French Edition)
SLR-Bench Versions:
SLR-Bench-French is the French-language pendant of the original SLR-Bench dataset.
It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into French.
This enables systematic evaluation and training of Large Language Models (LLMs) in logical reasoning in French… See the full description on the dataset page: https://huggingface.co/datasets/ahmad21omar/SLR-Bench-French.french-moore-parallel
French → Mooré (Mossi) Parallel Corpus
Machine-translated parallel sentences from French (fr) to Mooré / Mossi (mos), produced by a public-web crawl + filtering + Glosbe translation pipeline.
Snapshot
Field
Value
Validated pairs
3,000,040
Source language
French
Target language
Mooré (Mossi)
Translator
Glosbe public MT
Export date
2026-08-14
Schema
Column
Type
Description
id
string (UUID)
Pair identifier… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/french-moore-parallel.french_financial_news
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/arcticgiant/french-financial-news
Context
This dataset contains around 41 500 french news from 11/2018 to 03/2021 scraped on a famous financial media website.
For ease of use I’v add English translation (Helsinki-NLP/opus-mt-fr-en) and sentiment analysis (VADER)
Analysis
The picture below show the effect of covid crisis on news sentiment (Purple) and CAC40 (Blue).
We see clearly a link between the news sentiment… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/french_financial_news.
