datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Lucie-Training-Dataset
Lucie Training Dataset Card
The Lucie Training Dataset is a curated collection of text data
in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers,
digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages.
The Lucie Training Dataset was used to pretrain Lucie-7B,
a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.Luciole-Training-Dataset
Data card for The Luciole Training Dataset
Table of Contents
Dataset Description
Curation Rationale
Web Data Opt-Outs
Personal and Sensitive Information (PII)
Bias, Risks, and Limitations
Recommendations
Sample Metadata
Downloading the Data
Sample Use in Python
Accessing the English Web Data and OpenMathInstruct-1
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.Luciole-PostTraining-Dataset-1.1
Table of Contents
Dataset Description
Curation Rationale
Bias, Risks, and Limitations
Data Subsets
Sample Metadata
Downloading the Data
Available Configurations
Loading Examples
Accessing Data Through the Directory Hierarchy
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole-PostTraining-Dataset-1.1 is a curated collection of open, instruction-style text data designed for language model post training. It includes a mixture of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-PostTraining-Dataset-1.1.Brevets-Francais-1981-2026-Clean
🇫🇷 Brevets français 1981–2026 — Clean 🇫🇷
Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet).
Format : Parquet, prêt pour chargement streaming / distribué.
Source
Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI).
Génération
Entrée :… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/Brevets-Francais-1981-2026-Clean.wikipedia
Plain text of Wikipedia
Dataset Description
Size
Example use (python)
Data fields
Notes on data formatting
License
Aknowledgements
Citation
Dataset Description
This dataset is a plain text version of pages from wikipedia.org spaces for several languages
(English,
German,
French,
Spanish,
Italian).
The text is without HTML tags nor wiki templates.
It just includes markdown syntax for headers, lists and tables.
See Notes on data formatting for more details.
It was… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikipedia.Brevets-Francais-2000-2026-RawBrevets-Francais-1981-2026-Raw
🇫🇷 Brevets français 1981–2026 — Raw 🇫🇷
Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet).
Format : Parquet
Source
Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI).
Génération
Entrée :
822 310 fichiers XML
Sortie :
822 310 lignes… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/Brevets-Francais-1981-2026-Raw.wikimedia
Dataset Card
This dataset is a curated collection of Wikimedia pages in markdown format,
compiled from various Wikimedia projects across multiple languages.
Covered Wikimedia Projects:
wikipedia
wikibooks
wikinews
wikiquote
wikisource
wikiversity
wikivoyage
wiktionary
Supported Languages:
ar (Arabic)
br (Breton)
ca (Catalan)
co (Corsican)
de (German)
en (English)
es (Spanish)
eu (Basque)
fr (French)
frp (Arpitan)
it (Italian)
nl (Dutch)
oc (Occitan)
pcd (Picard)
pt (Portuguese)… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikimedia.Nemotron-Personas-France
Nemotron-Personas-France
Une approche d'IA composée pour des personas ancrés dans des distributions réelles
A compound AI approach to personas grounded in real-world distributions
Vue d'ensemble du jeu de données (Dataset Overview)
Nemotron-Personas-France est un jeu de données en libre accès (CC BY 4.0) composé de personas générés de manière synthétique. Ce jeu de données s'appuie sur les distributions démographiques, géographiques et de traits de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-France.horse-racing-france
French Horse Racing Dataset
Structured dataset of French horse racing covering 2014-01-01 → 2026-05-05,
restricted to French meetings.
This repository is provided "as is" under an other license.
Overview
Longitudinal collection of French race meetings with full race, runner, and
payout records. The data is split into four normalized tables, each exposed as
an independent Hub configuration.
Configuration
Description
Rows
reunions
Race meetings (one row… See the full description on the dataset page: https://huggingface.co/datasets/annaelmoussa/horse-racing-france.Brevets-Francais-2020-2026-RawRULER-luciole_tokenizer_128k-arab-regional_v2Brevets-Francais-2025-Claimswho-killed-jfkTranslation-Instruct
Corpus overview
Translation Instruct is a collection of parallel corpora for machine translation, formatted as instructions for the supervised fine-tuning of large language models. It currently contains two collections: Croissant Aligned Instruct and Europarl Aligned Instruct.
Croissant Aligned Instruct is an instruction-formatted version of the parallel French-English data in croissantllm/croissant_dataset_no_web_data
(subset: aligned_36b).
Europarl Aligned Instruct is an… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Translation-Instruct.wikisource
Plain text of Wikisource
Dataset Description
Size
Example use (python)
Data fields
Notes on data formatting
License
Aknowledgements
Citation
Dataset Description
This dataset is a plain text version of pages from wikisource.org in French language.
The text is without HTML tags nor wiki templates.
It just includes markdown syntax for headers, lists and tables.
See Notes on data formatting for more details.
It was created by LINAGORA and OpenLLM France
from the… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikisource.meteo-france-climatologie-quotidienne
Climatologie quotidienne Météo-France
Archive des observations climatologiques quotidiennes des stations
Météo-France, de 1786-05-01 à 2026-09-10,
en Parquet partitionné par département et par période.
Redistribution non officielle. Ce dépôt n'est pas publié par
Météo-France. Il redistribue des données publiques sous Licence Ouverte 2.0.
Source et attribution
Source : Météo-France, données publiques de climatologie de base.
Moissonnées depuis… See the full description on the dataset page: https://huggingface.co/datasets/cm3r/meteo-france-climatologie-quotidienne.Claire-Dialogue-English-0.1
Claire English Dialogue Dataset (CEDD) A collection of English dialogue transcripts
This is the first packaged version of the datasets used to train the english variants of the Claire family of large language models
(OpenLLM-France/Claire-7B-EN-0.1). (A related French dataset can be found here.)
The Claire English Dialogue Dataset (CEDD) is a collection of transcripts of English dialogues from various sources, including parliamentary proceedings, interviews, broadcast, meetings, and… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Claire-Dialogue-English-0.1.mantis_datasetfrancecrops
FranceCrops
FranceCrops is a crop-classification benchmark for French agricultural parcels observed with Sentinel-2 L2A time series meant to evaluate the representation learned by self-supervised or unsupervised methods. Each sample is one parcel represented by 100 sampled pixel time series. This release provides fixed supervised splits for downstream evaluation and frozen low-label subsets from 1 to 4,000 labels per class so methods can be compared under the same downstream… See the full description on the dataset page: https://huggingface.co/datasets/saget-antoine/francecrops.sg-ego
SG-Ego Annotations
Project Page: https://francescapistilli.github.io/GLEN.
Data annotation pipeline: https://github.com/francescapistilli/sg-ego
SG-Ego is a large scale annotation set extending Ego4D with spatio-temporal scene graphs, where relations triplets are consolidated over time into explicit time-evolving descriptions of the scene state.
SG-Ego is released as part of our paper "Learning to Evolve Scenes: Reasoning about Human Activities with Scene Graphs", which presents… See the full description on the dataset page: https://huggingface.co/datasets/francescapistilli/sg-ego.French-Patent-1981-2026-Clean
🇫🇷 Brevets français 1981–2026 — Clean 🇫🇷
Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet).
Format : Parquet, prêt pour chargement streaming / distribué.
Source
Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI).
Génération… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/French-Patent-1981-2026-Clean.Brevets-Francais-2024-Chunked
🇫🇷 Brevets français 2024 Chunké 🇫🇷
Dataset de brevets français publiés en 2024, extrait depuis les XML d’origine et chunké au niveau des balises <p> xml
Format : Parquet, prêt pour chargement streaming / distribué.
Source
Données issues de documents publics de brevets français (A1, 2024).Extraction, structuration et découpage réalisés de manière indépendante grace a un acces aux API/FTP PI (sur demande a l'inpi).
Génération
Entrée :
35 479 fichiers XML… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/Brevets-Francais-2024-Chunked.weather_france_radar_satellite_gs
This dataset is a ready-to-use fusion of multiple datasets for the France region, including:
Ground station data: Sourced from data.gouv.fr with a 1-hour resolution. It includes 7 weather KPIs: "RR1", "FF", "DD", "T", "U", "PMER", and "VV".
Radar imagery: Radar rainfall accumulation images from the Météo-France open data initiative.
Geospatial data: Land cover and ground height information from EarthEnv and OpenTopography.
Satellite imagery: Sourced from the EUMETSAT platform… See the full description on the dataset page: https://huggingface.co/datasets/meteolibre-dev/weather_france_radar_satellite_gs.Claire-Dialogue-French-0.1
Claire French Dialogue Dataset (CFDD) A collection of French dialogue transcripts and plays
This is the first packaged version of the datasets used to train the Claire family of large language models
(OpenLLM-France/Claire-7B-0.1).
The Claire French Dialogue Dataset (CFDD) is a collection of theater plays and transcripts of real French dialogues from various sources, including parliamentary proceedings, interviews, debates, meetings, and free conversations.
Each dialogue is split… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Claire-Dialogue-French-0.1.French-Patents-2020-2026-Raw
🇫🇷 Brevets français 2020–2026 — RAW 🇫🇷
Dataset de brevets français publiés entre 2020 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet).
Format : Parquet, prêt pour chargement streaming / distribué.
Source
Données issues de documents publics de brevets français (A1).Extraction réalisé de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI).
Génération
468 000 fichiers XML… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/French-Patents-2020-2026-Raw.Nemotron-Personas-France-Qwen3-0.6B-embedding
Nemotron-Personas-France-Qwen3-0.6B-embedding
Embeddings for nvidia/Nemotron-Personas-France computed with Qwen/Qwen3-Embedding-0.6B.
Details
Source dataset: nvidia/Nemotron-Personas-France
Embedding model: Qwen/Qwen3-Embedding-0.6B
Embedding dimension: 1024
Number of rows: 1000000 (first 1M rows of the 7M-row source dataset)
Columns embedded (in dataset order): professional_persona, sports_persona, arts_persona, travel_persona, culinary_persona, persona… See the full description on the dataset page: https://huggingface.co/datasets/tantara/Nemotron-Personas-France-Qwen3-0.6B-embedding.comp-mechEach record in the dataset contains the following fields:
target_new: the counterfactual term
target_true: the actual term
subject: the topic of the prompt
base_prompt: the foundational prompt
prompt: the modified prompt incorporating the counterfactual change
template: the sentence structure using the counterfactual wording.
ipfs_france_laws_ir
France legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_france_laws (revision ``) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of France prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was invented.
Primary key: entry_cid (CIDv1 raw… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_france_laws_ir.Luciole_RAG
Dataset overview
Luciole RAG is a supervised fine-tuning dataset for retrieval-augmented generation, built to train the Luciole models. Each example is a chat conversation where the assistant answers a question using only a set of retrieved document chunks given in the system prompt, quotes and cites its sources, and declines to answer when the documents do not contain the answer.
It contains two subsets derived from existing question-answering benchmarks:
Config
Source… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole_RAG.
