datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Lucie-Training-Dataset
Lucie Training Dataset Card
The Lucie Training Dataset is a curated collection of text data
in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers,
digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages.
The Lucie Training Dataset was used to pretrain Lucie-7B,
a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.France_Government_ConversationsSeasonBench-EA-METEO_FRANCE-PressureThe dataset is a subset for SeasonBench-EA Benchmark, which contains the ensemble forecasts from Meteo France on pressure levels. All data are downloaded from the Copernicus Climate Data Store (https://cds.climate.copernicus.eu/) and reorganized for benchmark construction. Please ensure compliance with the CDS Licence when using or redistributing the data.
Luciole-Training-Dataset
Data card for The Luciole Training Dataset
Table of Contents
Dataset Description
Curation Rationale
Web Data Opt-Outs
Personal and Sensitive Information (PII)
Bias, Risks, and Limitations
Recommendations
Sample Metadata
Downloading the Data
Sample Use in Python
Accessing the English Web Data and OpenMathInstruct-1
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.Luciole-PostTraining-Dataset-1.1
Table of Contents
Dataset Description
Curation Rationale
Bias, Risks, and Limitations
Data Subsets
Sample Metadata
Downloading the Data
Available Configurations
Loading Examples
Accessing Data Through the Directory Hierarchy
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole-PostTraining-Dataset-1.1 is a curated collection of open, instruction-style text data designed for language model post training. It includes a mixture of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-PostTraining-Dataset-1.1.Brevets-Francais-1981-2026-Clean
🇫🇷 Brevets français 1981–2026 — Clean 🇫🇷
Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet).
Format : Parquet, prêt pour chargement streaming / distribué.
Source
Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI).
Génération
Entrée :… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/Brevets-Francais-1981-2026-Clean.wikipedia
Plain text of Wikipedia
Dataset Description
Size
Example use (python)
Data fields
Notes on data formatting
License
Aknowledgements
Citation
Dataset Description
This dataset is a plain text version of pages from wikipedia.org spaces for several languages
(English,
German,
French,
Spanish,
Italian).
The text is without HTML tags nor wiki templates.
It just includes markdown syntax for headers, lists and tables.
See Notes on data formatting for more details.
It was… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikipedia.Brevets-Francais-2000-2026-RawBrevets-Francais-1981-2026-Raw
🇫🇷 Brevets français 1981–2026 — Raw 🇫🇷
Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet).
Format : Parquet
Source
Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI).
Génération
Entrée :
822 310 fichiers XML
Sortie :
822 310 lignes… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/Brevets-Francais-1981-2026-Raw.wikimedia
Dataset Card
This dataset is a curated collection of Wikimedia pages in markdown format,
compiled from various Wikimedia projects across multiple languages.
Covered Wikimedia Projects:
wikipedia
wikibooks
wikinews
wikiquote
wikisource
wikiversity
wikivoyage
wiktionary
Supported Languages:
ar (Arabic)
br (Breton)
ca (Catalan)
co (Corsican)
de (German)
en (English)
es (Spanish)
eu (Basque)
fr (French)
frp (Arpitan)
it (Italian)
nl (Dutch)
oc (Occitan)
pcd (Picard)
pt (Portuguese)… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikimedia.Nemotron-Personas-France
Nemotron-Personas-France
Une approche d'IA composée pour des personas ancrés dans des distributions réelles
A compound AI approach to personas grounded in real-world distributions
Vue d'ensemble du jeu de données (Dataset Overview)
Nemotron-Personas-France est un jeu de données en libre accès (CC BY 4.0) composé de personas générés de manière synthétique. Ce jeu de données s'appuie sur les distributions démographiques, géographiques et de traits de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-France.horse-racing-france
French Horse Racing Dataset
Structured dataset of French horse racing covering 2014-01-01 → 2026-05-05,
restricted to French meetings.
This repository is provided "as is" under an other license.
Overview
Longitudinal collection of French race meetings with full race, runner, and
payout records. The data is split into four normalized tables, each exposed as
an independent Hub configuration.
Configuration
Description
Rows
reunions
Race meetings (one row… See the full description on the dataset page: https://huggingface.co/datasets/annaelmoussa/horse-racing-france.Brevets-Francais-2020-2026-Rawnave-hidalgo-dataset
Nave Hidalgo — dataset completo del vuelo de inspeccion
Capturas originales, sin procesar, del vuelo de inspeccion de una nave industrial y su
estacionamiento (27 jul 2026) con un DJI Matrice 4T.
Fotografias RGB (*_V.JPG)
3,216 · 4032x3024 · ~10.9 GB
Fotografias termicas (*_T.JPG)
3,215 · 1280x1024 · ~5.8 GB
Coordenadas del sitio
20.054569, -99.311673
Emparejado RGB <-> termico
Cada toma existe como par. Comparten el prefijo completo (marca… See the full description on the dataset page: https://huggingface.co/datasets/Francel12/nave-hidalgo-dataset.france-climate-risk-drias-tracc-2023
France Climate Risk DRIAS TRACC-2023
Ce dataset regroupe des fichiers NetCDF issus de DRIAS TRACC-2023 pour la France métropolitaine, organisés pour un usage plus simple.
Le dépôt contient des indicateurs climatiques :
pour les niveaux de réchauffement +2.0 °C, +2.7 °C et +4.0 °C ;
en produits multi-modèles ;
en produits par modèle ;
pour les valeurs absolues ;
ainsi que pour la période de référence historique.
Structure du dépôt
Le dépôt est organisé par niveau de… See the full description on the dataset page: https://huggingface.co/datasets/saadtaleb/france-climate-risk-drias-tracc-2023.RULER-luciole_tokenizer_128k-arab-regional_v2Brevets-Francais-2025-ClaimsSeasonBench-EA-METEO_FRANCE-SingleLevelThe dataset is a subset for SeasonBench-EA Benchmark, which contains the ensemble forecasts from Meteo France on single level. All data are downloaded from the Copernicus Climate Data Store (https://cds.climate.copernicus.eu/) and reorganized for benchmark construction. Please ensure compliance with the CDS Licence when using or redistributing the data.
who-killed-jfkTranslation-Instruct
Corpus overview
Translation Instruct is a collection of parallel corpora for machine translation, formatted as instructions for the supervised fine-tuning of large language models. It currently contains two collections: Croissant Aligned Instruct and Europarl Aligned Instruct.
Croissant Aligned Instruct is an instruction-formatted version of the parallel French-English data in croissantllm/croissant_dataset_no_web_data
(subset: aligned_36b).
Europarl Aligned Instruct is an… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Translation-Instruct.wikisource
Plain text of Wikisource
Dataset Description
Size
Example use (python)
Data fields
Notes on data formatting
License
Aknowledgements
Citation
Dataset Description
This dataset is a plain text version of pages from wikisource.org in French language.
The text is without HTML tags nor wiki templates.
It just includes markdown syntax for headers, lists and tables.
See Notes on data formatting for more details.
It was created by LINAGORA and OpenLLM France
from the… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikisource.meteo-france-climatologie-quotidienne
Climatologie quotidienne Météo-France
Archive des observations climatologiques quotidiennes des stations
Météo-France, de 1786-05-01 à 2026-09-10,
en Parquet partitionné par département et par période.
Redistribution non officielle. Ce dépôt n'est pas publié par
Météo-France. Il redistribue des données publiques sous Licence Ouverte 2.0.
Source et attribution
Source : Météo-France, données publiques de climatologie de base.
Moissonnées depuis… See the full description on the dataset page: https://huggingface.co/datasets/cm3r/meteo-france-climatologie-quotidienne.tl-climate-projectionscrosswalks-france-samvehicles-q0x2v
Dataset Card for vehicles-q0x2v
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
vehicles-q0x2v
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/vehicles-q0x2v.axial-mri
Dataset Card for axial-mri
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
axial-mri
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=640x640… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/axial-mri.Claire-Dialogue-English-0.1
Claire English Dialogue Dataset (CEDD) A collection of English dialogue transcripts
This is the first packaged version of the datasets used to train the english variants of the Claire family of large language models
(OpenLLM-France/Claire-7B-EN-0.1). (A related French dataset can be found here.)
The Claire English Dialogue Dataset (CEDD) is a collection of transcripts of English dialogues from various sources, including parliamentary proceedings, interviews, broadcast, meetings, and… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Claire-Dialogue-English-0.1.mantis_datasetpedestrian-crosswalks-france-v2insects-mytwu
Dataset Card for insects-mytwu
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
insects-mytwu
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/insects-mytwu.
