datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
french_CEFRddxplus-french
Dataset Description
We are releasing under the CC-BY licence a new large-scale dataset for Automatic Symptom Detection (ASD) and Automatic Diagnosis (AD) systems in the medical domain. The dataset contains patients synthesized using a proprietary medical knowledge base and a commercial rule-based AD system. Patients in the dataset are characterized by their socio-demographic data, a pathology they are suffering from, a set of symptoms and antecedents related to this pathology, and a… See the full description on the dataset page: https://huggingface.co/datasets/aai530-group6/ddxplus-french.cold-french-law
Collaborative Open Legal Data (COLD) - French Law
COLD French Law is a dataset containing over 800 000 french law articles, filtered and extracted from France's LEGI dataset and formatted as a single CSV file.
This dataset focuses on articles (codes, lois, décrets, arrêtés ...) identified as currently applicable french law.
A large portion of this dataset comes with machine-generated english translations, provided by Casetext, Part of Thomson Reuters using OpenAI's GPT-4.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-french-law.diverse_french_newsfrench-triviaFrenchHateSpeechSuperset
FrenchHateSpeechSuperset
This dataset is a superset of multiple datasets including hate speech, harasment, sexist, racist, etc...messages from various platforms.
Included datasets :
MLMA dataset
CAA dataset
FTR dataset
"An Annotated Corpus for Sexism Detection in French Tweets" dataset
UC-Berkeley-Measuring-Hate-Speech dataset (translated from english*)
References
@inproceedings{chiril2020annotated,
title={An Annotated Corpus for Sexism Detection in French Tweets}… See the full description on the dataset page: https://huggingface.co/datasets/Poulpidot/FrenchHateSpeechSuperset.rte3-french
Dataset Card for French RTE-3
Dataset Summary
The RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge.
Like its English counterpart, the French RTE-3 dataset is composed of a development set and a test set, each containing 800 T/H pairs.
All T/H pairs were manually translated into French and proofread.
It is annotated for a 3-way task.
Please refer to this repository, if you want to use this French version… See the full description on the dataset page: https://huggingface.co/datasets/maximoss/rte3-french.hatecheck-french
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-french.Wolof-to-French_Translation-Dataset
Dataset Wolof ↔ Français
🧩 Présentation
Ce dataset contient plus de 30 000 paires phrase Wolof – phrase Française.Chaque ligne est structurée comme suit :
Wolof (input)
Français (target)
Phrase en Wolof
Phrase correspondante en Français
Il a été conçu pour la traduction automatique et les tâches de NLP impliquant le Wolof et le Français.
📚 Provenance et nettoyage
Le dataset a été créé en compilant différentes sources accessibles… See the full description on the dataset page: https://huggingface.co/datasets/MaroneAI/Wolof-to-French_Translation-Dataset.spelling-correction-french-news
Spelling correction dataset (French)
This dataset is generated by transforming/corrupting sentences of a French news corpus
provided by the University of Leipzig.
The following transformations are applied to words in the sentences:
concatenation of pairs of words
swapping of neighboring letters in words
insertion
deletion
replacement (by neighboring characters in AZERTY keyboard)
Generation
./scripts/get_data.py -t news -y 2023 -s 10K
./scripts/generate_dataset.py… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/spelling-correction-french-news.french_engfrench_narrativeqa
Description
Dataframe containing 143 French books in txt format.More precisely :
the texte column contains the texts
the titre column contains the book title
the auteur column contains the author's name and dates of birth and death (if you want to filter the texts to keep only those from the given century to the present day)
the question column contains a single question asked about the associated text
the answers column contains one or more answers to the question (= if several… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_narrativeqa.erudit-french-philosophy
Dataset Card for Dataset Name
Dataset Description
Dataset Summary
This dataset contains all french philosophy that has been published on erudit.org. It has been generated using a Bs4 web parser that you can find in this repo: https://github.com/MFGiguere/french-philosophy-generator.
Supported Tasks and Leaderboards
This dataset could be useful for this (non-exhaustive) set of tasks: detect if a text is philosophical or not, generate philosophical… See the full description on the dataset page: https://huggingface.co/datasets/mfgiguere/erudit-french-philosophy.chatgpt-prompts-Frenchfrench_rap_songsfrench-brand-content-benchmark-2026
French Brand Content Benchmark 2026
Publisher: Big NeuronsWebsite: https://www.bigneurons.comEnglish version: https://www.bigneurons.com/enContact: brief@bigneurons.comLicense: CC BY 4.0Last updated: March 2026DOI: 10.5281/zenodo.18927033tags:
brand-content
marketing
france
benchmark
acquisition
geo
What is this dataset?
The French Brand Content Benchmark 2026 is the first publicly available benchmark of brand content performance metrics for French SMEs and… See the full description on the dataset page: https://huggingface.co/datasets/BigNeurons/french-brand-content-benchmark-2026.french_financial_news
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/arcticgiant/french-financial-news
Context
This dataset contains around 41 500 french news from 11/2018 to 03/2021 scraped on a famous financial media website.
For ease of use I’v add English translation (Helsinki-NLP/opus-mt-fr-en) and sentiment analysis (VADER)
Analysis
The picture below show the effect of covid crisis on news sentiment (Purple) and CAC40 (Blue).
We see clearly a link between the news sentiment… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/french_financial_news.French_Doctoral_Theses
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/antoinebourgois2/french-doctoral-thesis
Description
All french doctoral thesis metatdata scrapped from https://www.theses.fr
The dataset contains :
URL
Thesis title
Short description
Author name
Thesis director(s) informations
Research domain
Status ( defended / in preparation)
french_first_names_insee_2024
French First Names from Death Records (1970-2024)
This dataset contains French first names extracted from death records provided by INSEE (French National Institute of Statistics and Economic Studies) covering the period from 1970 to September 2024.
Dataset Description
Data Source
The data is sourced from INSEE's death records database. It includes first names of deceased individuals in France, providing valuable insights into naming patterns across different… See the full description on the dataset page: https://huggingface.co/datasets/eltorio/french_first_names_insee_2024.french_books
Description
Dataframe containing 2075 French books in txt format (= the ~2600 French books present in gutenberg from which all books by authors present in the french_books_summuries dataset have been removed to avoid any leaks).More precisely :
the texte column contains the texts
the titre column contains the book title
the auteur column contains the author's name and dates of birth and death (if you want to filter the texts to keep only those from the given century to the present… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_books.Pedale-FrenchTextCorpus
Pedale-FrenchTextCorpus
tags: classification, linguistics, French
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'Pedale-FrenchTextCorpus' is a collection of French text excerpts from articles, blog posts, and forum discussions related to bicycle pedals ('pedales'). The dataset is intended for machine learning practitioners who are working on text classification tasks that focus on French-language content, with an emphasis on… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/Pedale-FrenchTextCorpus.Kabyle-French
French - Kabyle (Tatoeba)
This dataset contains translation pairs for French (fr) and Kabyle (kab). The data was collected from the Tatoeba Project, a free collaborative online database of example sentences.
⚠️ Important Note on Quality
Disclaimer: This dataset has been exported automatically and has not been manually verified. While Tatoeba relies on community contributions, errors or inconsistencies in translation pairs may exist. Use with appropriate caution.
French-reviews-with-prediction-of-their-feelings
🇫🇷 French Sentiment Dataset - Multisource (Auto-labeled)
This dataset contains 960,000 French-language text reviews, automatically labeled with sentiment classes: positive, neutral, or negative, using the TextBlob-FR library.
It is suitable for training or evaluating sentiment classification models in French, as well as general-purpose NLP research on French texts.
⚙️ Labeling method
Sentiment labels were automatically generated based on polarity scores using the… See the full description on the dataset page: https://huggingface.co/datasets/gaouehalim/French-reviews-with-prediction-of-their-feelings.french-local-authorities-payment-delays
Payment delays of French local authorities, 2024 and 2025
How long French local authorities take to pay their suppliers, budget by budget.
182 763 records covering two fiscal years, with the average annual payment delay
of each authority and whether it meets the 30-day statutory limit.
Open public data
This dataset is derived from open public data published by the French
Direction générale des finances publiques (DGFiP) on
data.gouv.fr, under the
Open Licence 2.0.… See the full description on the dataset page: https://huggingface.co/datasets/freginer/french-local-authorities-payment-delays.frenchnews-7
FrenchNews-7
FrenchNews-7 is a cross-publisher French news editorial desk classification benchmark. This public release is a manifest-only artifact for reproducibility and benchmarking: it exposes URL-level and metadata-level information for 87,637 labeled articles from 13 publishers without redistributing article text.
The target variable is the publisher's editorial routing decision — which desk a newsroom assigned an article to — rather than an annotator's perceived topic.… See the full description on the dataset page: https://huggingface.co/datasets/LeFrenchNewsLab/frenchnews-7.French_ofFrench_GEC
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/isakbiderre/french-gec-dataset
Context
Wikipedia is a free encyclopedia where everyone can contribute and modify, delete, or add text to the articles.
Because of this, every day there is newly created text and, most importantly, new corrections made to preexisting sentences.
The idea is to find the corrections made to these sentences and create a dataset with X,y sentence pairs.
The data
This dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/French_GEC.French_Grammar_Explanations
This dataset contains 1500+ French grammar explanations. It's the one I used to train my finetuned LLM called FrenchLlama-3.2-1B-Instruct.
You can use this dataset for your own training purposes & find the aforementioned model on my HuggingFace profile.
French_Wolof_Various_Parallel_Corpusfrench-hate-speech-superset
French Hate Speech Superset
This dataset is a superset (N=18,071) of posts annotated as hateful or not. It results from the preprocessing and merge of all available French hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior, that… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/french-hate-speech-superset.
