France
Datasets
All datasets matching “France”Lucie-Training-Dataset
Lucie Training Dataset Card
The Lucie Training Dataset is a curated collection of text data
in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers,
digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages.
The Lucie Training Dataset was used to pretrain Lucie-7B,
a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.France_Government_ConversationsSeasonBench-EA-METEO_FRANCE-PressureThe dataset is a subset for SeasonBench-EA Benchmark, which contains the ensemble forecasts from Meteo France on pressure levels. All data are downloaded from the Copernicus Climate Data Store (https://cds.climate.copernicus.eu/) and reorganized for benchmark construction. Please ensure compliance with the CDS Licence when using or redistributing the data.
Luciole-Training-Dataset
Data card for The Luciole Training Dataset
Table of Contents
Dataset Description
Curation Rationale
Web Data Opt-Outs
Personal and Sensitive Information (PII)
Bias, Risks, and Limitations
Recommendations
Sample Metadata
Downloading the Data
Sample Use in Python
Accessing the English Web Data and OpenMathInstruct-1
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.Luciole-PostTraining-Dataset-1.1
Table of Contents
Dataset Description
Curation Rationale
Bias, Risks, and Limitations
Data Subsets
Sample Metadata
Downloading the Data
Available Configurations
Loading Examples
Accessing Data Through the Directory Hierarchy
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole-PostTraining-Dataset-1.1 is a curated collection of open, instruction-style text data designed for language model post training. It includes a mixture of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-PostTraining-Dataset-1.1.Brevets-Francais-1981-2026-Clean
🇫🇷 Brevets français 1981–2026 — Clean 🇫🇷
Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet).
Format : Parquet, prêt pour chargement streaming / distribué.
Source
Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI).
Génération
Entrée :… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/Brevets-Francais-1981-2026-Clean.
