datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
French-PD-Newspapers
🇫🇷 French Public Domain Newspapers 🇫🇷
French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.French-PD-Books
🇫🇷 French Public Domain Books 🇫🇷
French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.french_instruct
🧑🏫 French Instruct
The French Instruct dataset is a collection of instructions with their corresponding answers (sometimes multi-turn conversations) entirely in French. The dataset is also available on GitHub.
📊 Overview
The dataset is composed of 276K conversations between a user and an assistant for a total of approximately 85M tokens.
I also added annotations for each document to indicate if it was generated or written by a human, the style of… See the full description on the dataset page: https://huggingface.co/datasets/angeluriot/french_instruct.eighteenth_century_french_novels
General information
This dataset contains 12 Mio Token of Literary French prose 1751-1800 in plain text format, built within the project 'Mining and Modeling Text' (2019-2023) at Trier University.
For the dataset in XML/TEI see the GitHub repository of the project.
Collection de romans français du dix-huitième siècle (1751-1800) / Collection of Eighteenth-Century French Novels (1751-1800)
This collection of Eighteenth-Century French Novels contains 200 digital… See the full description on the dataset page: https://huggingface.co/datasets/roettger/eighteenth_century_french_novels.Argimi-Legal-French-Jurisprudence
The ArGiMi French Jurisprudence Dataset
This dataset contains a comprehensive collection of French case law, sourced from the official archives of French jurisprudence. It is divided into three distinct subdivisions: Constitutional ("constit"), Administrative ("cetat"), and Judiciary ("juri").
This dataset was created for the ArGiMi project, an open-source initiative dedicated to promoting open data and knowledge sharing. The project is a collaborative effort between Giskard… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Legal-French-Jurisprudence.French_Wikipedia_articlesDump of 2023-08-20 of all french article in wikipedia
https://dumps.wikimedia.org/frwiki/20230820/frwiki-20230820-pages-articles.xml.bz2
french-fiction-16-18th-century
French Fiction of the 16th–18th Centuries
A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model.
The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction.
Structure
Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.Claire-Dialogue-French-0.1
Claire French Dialogue Dataset (CFDD) A collection of French dialogue transcripts and plays
This is the first packaged version of the datasets used to train the Claire family of large language models
(OpenLLM-France/Claire-7B-0.1).
The Claire French Dialogue Dataset (CFDD) is a collection of theater plays and transcripts of real French dialogues from various sources, including parliamentary proceedings, interviews, debates, meetings, and free conversations.
Each dialogue is split… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Claire-Dialogue-French-0.1.English_French_Songs_Lyrics_Translation_Original
Original Songs Lyrics with French Translation
Dataset Summary
Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French.
Details of the number of songs by language of origin can be found in the table below:
Original language
Number of songs
en
75786
fr
18486
es
1743
it
803
de
691
sw
529
ko
193
id
169
pt
142
no
122
fi
113
sv
70
hr
53
so
43
ca
41
tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.task815_pawsx_japanese_french_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task815_pawsx_japanese_french_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task815_pawsx_japanese_french_translation.french-dev-questions-5k
French Developer Questions 5K
5,000 software-engineering questions written in French, generated by 22 open
models across a combinatorial seed grid, then filtered one by one by an LLM judge.
No answers — this is a prompt corpus, meant to be the input side of a
distillation or SFT pipeline.
French technical data is scarce. Most French datasets on the Hub are literary,
legal or journalistic corpora; most developer-question datasets are English-only.
This one sits in the… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/french-dev-questions-5k.banque-fr-2311
Dataset Card for "banque fr websites - 2311"
Dataset extracted from public websites by wordslab-webscraper in 2311:
domain: banque
language: fr
license: Apache 2.0
Dataset Sources
wordslab-webscraper follows the industry best practices for polite web scraping:
clearly identifies itself as a known text indexing bot: "bingbot"
doesn't try to hide the user IP address behind proxies
doesn't try to circumvent bots protection solutions
waits for a minimum delay between two… See the full description on the dataset page: https://huggingface.co/datasets/frenchtext/banque-fr-2311.task778_pawsx_english_french_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task778_pawsx_english_french_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task778_pawsx_english_french_translation.french_orca_dpo_pairs
Dataset Card for french_orca_dpo_pairs
This dataset offers a french translation of the 12k DPO Intel/orca_dpo_pairs pairs made from Open-Orca/OpenOrca.
Dataset Card Contact
ntnq
bank-en-2401
Dataset Card for "bank en websites - 2401"
Dataset extracted from public websites by wordslab-webscraper in 2401:
domain: bank
language: en
license: Apache 2.0
Dataset Sources
wordslab-webscraper follows the industry best practices for polite web scraping:
clearly identifies itself as a known text indexing bot: "bingbot"
doesn't try to hide the user IP address behind proxies
doesn't try to circumvent bots protection solutions
waits for a minimum delay between two pages… See the full description on the dataset page: https://huggingface.co/datasets/frenchtext/bank-en-2401.jfv-french-style-conditioning-dataset-v1.0
JFV French Paired Style-Conditioning Dataset
At a Glance
Item
Value
Language
French
Source
Single-author blog corpus, 2005–2025
Public release
v1.0
Aligned units in public release
1,484
Texts in public aligned release
7,420
Original experiment
1,492 aligned units / 7,460 texts
Generated conditions
Ministral baseline; profile; profile + five-shot examples
Primary use
Paired study of stylistic conditioning and evaluation-metric validity… See the full description on the dataset page: https://huggingface.co/datasets/PeggyVallin/jfv-french-style-conditioning-dataset-v1.0.frenchPARAPHRASE
Dataset information
Dataset concatenating Paraphases datasets available in French and open-source.There are a total of 254,513 rows, of which 251,753 are for training, 1,857 for validation and 903 for testing.
Usage
from datasets import load_dataset
dataset = load_dataset("CATIE-AQ/frenchPARAPHRASE")
Dataset
Details of rows
Dataset Original
Splits
Note
Helsinki-NLP/tatoeba_mt
2,117 train / 999 validation
We only keep the French split… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/frenchPARAPHRASE.French_Grammar_Explanations
This dataset contains 1500+ French grammar explanations. It's the one I used to train my finetuned LLM called FrenchLlama-3.2-1B-Instruct.
You can use this dataset for your own training purposes & find the aforementioned model on my HuggingFace profile.
french-senate-session-reports
🏛️ French Senate Session Reports Dataset
A dataset of parliamentary debates and sessions reports from the French Senate.508,647,861 tokens of high-quality French text transcribed manually from Senate Sessions
Description
This dataset consists of all session reports from the French Senate debates, crawled from the official website senat.fr. It provides high-quality text data of parliamentary discussions, covering a wide range of political, economic, and social topics… See the full description on the dataset page: https://huggingface.co/datasets/TheJeanneCompany/french-senate-session-reports.task483_cls_french_dvd_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task483_cls_french_dvd_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task483_cls_french_dvd_classification.mistral-legal-french-dataset
Mistral Legal French Dataset
A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy.
📋 Table of Contents
Overview
Dataset Composition
Methodology
1. Chain-of-Thought Generation
2. LegalKit Extraction
3. Curriculum Learning Fusion
Data Format
Quality Metrics
Usage
Citations
License
🎯 Overview
This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two… See the full description on the dataset page: https://huggingface.co/datasets/davidpistori/mistral-legal-french-dataset.bank-de-2401
Dataset Card for "bank de websites - 2401"
Dataset extracted from public websites by wordslab-webscraper in 2401:
domain: bank
language: de
license: Apache 2.0
Dataset Sources
wordslab-webscraper follows the industry best practices for polite web scraping:
clearly identifies itself as a known text indexing bot: "bingbot"
doesn't try to hide the user IP address behind proxies
doesn't try to circumvent bots protection solutions
waits for a minimum delay between two pages… See the full description on the dataset page: https://huggingface.co/datasets/frenchtext/bank-de-2401.Alpaca_french_mixtral
Dataset Card for Alpaca_french_mixtral
This dataset was made by reusing the french alpaca instruction with Mixtral-8x7B-Instruct to make the output open-source.
Dataset Card Contact
robinjo
french_qa
Wiki-FR-QA : Dataset de Questions-Réponses en Français
Description
Dataset de question-answering en français généré automatiquement à partir d'articles
Wikipedia FR. Les questions sont produites par Qwen3.5-4B et filtrées pour la qualité.
Structure
Champ
Description
id
Identifiant unique (SHA-256 tronqué)
context
Section Wikipedia servant de contexte
question
Question en français
answer
Réponse en une phrase
article_title
Article source… See the full description on the dataset page: https://huggingface.co/datasets/CGCTG/french_qa.task484_cls_french_music_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task484_cls_french_music_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task484_cls_french_music_classification.FairytaleQA-translated-french
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the French machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-french.french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review
french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review
Summary
french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 347,688 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset french_book_reviews.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review.squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question
squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question
Summary
squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question is a subset of the Dataset of French Prompts (DFP).It contains 1,271,928 rows that can be used for a context-generation (with answer and question) task.The original data (without prompts) comes from the dataset pragnakalp/squad_v2_french_translated and was augmented by questions in SQUAD 2.0 format in the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question.mistral-legal-french-dataset
Mistral Legal French Dataset
A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy.
📋 Table of Contents
Overview
Dataset Composition
Methodology
1. Chain-of-Thought Generation
2. LegalKit Extraction
3. Curriculum Learning Fusion
Data Format
Quality Metrics
Usage
Citations
License
🎯 Overview
This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two complementary… See the full description on the dataset page: https://huggingface.co/datasets/VinceGx33/mistral-legal-french-dataset.french-corpus-llm-sample
French Corpus LLM — Sample 500 (v1.4.0)
FINALEADS LLC builds compliance-ready training datasets for French regulated industries. We turn 2.66 billion tokens of French finance, regulatory, and economic open data into audit-trailed, pseudonymized, AI Act Article 10-documented shares — so foundation model and regtech teams can ship into European enterprises without a data-lineage gap.
This is a public sample of 500 stratified documents drawn from the French Premium Web Corpus v1.4.0… See the full description on the dataset page: https://huggingface.co/datasets/finaleads/french-corpus-llm-sample.
