CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenLLM-France /Lucie-Training-Dataset Lucie Training Dataset Card The Lucie Training Dataset is a curated collection of text data in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers, digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages. The Lucie Training Dataset was used to pretrain Lucie-7B, a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.texttext-generation10B<n<100B39 likes30k downloads1y agoHugging Face02OpenLLM-France /Luciole-Training-Dataset Data card for The Luciole Training Dataset Table of Contents Dataset Description Curation Rationale Web Data Opt-Outs Personal and Sensitive Information (PII) Bias, Risks, and Limitations Recommendations Sample Metadata Downloading the Data Sample Use in Python Accessing the English Web Data and OpenMathInstruct-1 Details on Data Sources Citation Acknowledgements Contact Dataset Description The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.texttext-generation1B<n<10B15 likes1.8k downloads2mo agoHugging Face03OpenLLM-France /Luciole-PostTraining-Dataset-1.1 Table of Contents Dataset Description Curation Rationale Bias, Risks, and Limitations Data Subsets Sample Metadata Downloading the Data Available Configurations Loading Examples Accessing Data Through the Directory Hierarchy Details on Data Sources Citation Acknowledgements Contact Dataset Description The Luciole-PostTraining-Dataset-1.1 is a curated collection of open, instruction-style text data designed for language model post training. It includes a mixture of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-PostTraining-Dataset-1.1.text1M<n<10M5 likes1.8k downloads17h agoHugging Face04INPI-France /Brevets-Francais-1981-2026-Clean 🇫🇷 Brevets français 1981–2026 — Clean 🇫🇷 Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet). Format : Parquet, prêt pour chargement streaming / distribué. Source Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI). Génération Entrée :… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/Brevets-Francais-1981-2026-Clean.tabular100K<n<1M1 likes1.1k downloads8mo agoHugging Face05OpenLLM-France /wikipedia Plain text of Wikipedia Dataset Description Size Example use (python) Data fields Notes on data formatting License Aknowledgements Citation Dataset Description This dataset is a plain text version of pages from wikipedia.org spaces for several languages (English, German, French, Spanish, Italian). The text is without HTML tags nor wiki templates. It just includes markdown syntax for headers, lists and tables. See Notes on data formatting for more details. It was… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikipedia.texttext-generation10M<n<100M5 likes1.1k downloads2y agoHugging Face06INPI-France /Brevets-Francais-2000-2026-Rawtext100K<n<1M1 likes1k downloads8mo agoHugging Face07INPI-France /Brevets-Francais-1981-2026-Raw 🇫🇷 Brevets français 1981–2026 — Raw 🇫🇷 Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet). Format : Parquet Source Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI). Génération Entrée : 822 310 fichiers XML Sortie : 822 310 lignes… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/Brevets-Francais-1981-2026-Raw.text100K<n<1M3 likes923 downloads5mo agoHugging Face08OpenLLM-France /wikimedia Dataset Card This dataset is a curated collection of Wikimedia pages in markdown format, compiled from various Wikimedia projects across multiple languages. Covered Wikimedia Projects: wikipedia wikibooks wikinews wikiquote wikisource wikiversity wikivoyage wiktionary Supported Languages: ar (Arabic) br (Breton) ca (Catalan) co (Corsican) de (German) en (English) es (Spanish) eu (Basque) fr (French) frp (Arpitan) it (Italian) nl (Dutch) oc (Occitan) pcd (Picard) pt (Portuguese)… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikimedia.texttext-generation10M<n<100M3 likes856 downloads1y agoHugging Face09nvidia /Nemotron-Personas-France Nemotron-Personas-France Une approche d'IA composée pour des personas ancrés dans des distributions réelles A compound AI approach to personas grounded in real-world distributions Vue d'ensemble du jeu de données (Dataset Overview) Nemotron-Personas-France est un jeu de données en libre accès (CC BY 4.0) composé de personas générés de manière synthétique. Ce jeu de données s'appuie sur les distributions démographiques, géographiques et de traits de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-France.imagetext-generation1M<n<10M89 likes837 downloads6mo agoHugging Face10annaelmoussa /horse-racing-france French Horse Racing Dataset Structured dataset of French horse racing covering 2014-01-01 → 2026-05-05, restricted to French meetings. This repository is provided "as is" under an other license. Overview Longitudinal collection of French race meetings with full race, runner, and payout records. The data is split into four normalized tables, each exposed as an independent Hub configuration. Configuration Description Rows reunions Race meetings (one row… See the full description on the dataset page: https://huggingface.co/datasets/annaelmoussa/horse-racing-france.tabularother1M<n<10M1 likes574 downloads2mo agoHugging Face11INPI-France /Brevets-Francais-2020-2026-Rawtext100K<n<1M1 likes528 downloads8mo agoHugging Face12OpenLLM-France /RULER-luciole_tokenizer_128k-arab-regional_v2tabular10K<n<100K0 likes374 downloads10mo agoHugging Face13INPI-France /Brevets-Francais-2025-Claimstabular100K<n<1M1 likes373 downloads7mo agoHugging Face14Francesco /who-killed-jfkimage10K<n<100K2 likes323 downloads1y agoHugging Face15OpenLLM-France /Translation-Instruct Corpus overview Translation Instruct is a collection of parallel corpora for machine translation, formatted as instructions for the supervised fine-tuning of large language models. It currently contains two collections: Croissant Aligned Instruct and Europarl Aligned Instruct. Croissant Aligned Instruct is an instruction-formatted version of the parallel French-English data in croissantllm/croissant_dataset_no_web_data (subset: aligned_36b). Europarl Aligned Instruct is an… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Translation-Instruct.texttext-generation100K<n<1M5 likes295 downloads1y agoHugging Face16OpenLLM-France /wikisource Plain text of Wikisource Dataset Description Size Example use (python) Data fields Notes on data formatting License Aknowledgements Citation Dataset Description This dataset is a plain text version of pages from wikisource.org in French language. The text is without HTML tags nor wiki templates. It just includes markdown syntax for headers, lists and tables. See Notes on data formatting for more details. It was created by LINAGORA and OpenLLM France from the… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikisource.texttext-generation100K<n<1M4 likes287 downloads2y agoHugging Face17cm3r /meteo-france-climatologie-quotidienne Climatologie quotidienne Météo-France Archive des observations climatologiques quotidiennes des stations Météo-France, de 1786-05-01 à 2026-09-10, en Parquet partitionné par département et par période. Redistribution non officielle. Ce dépôt n'est pas publié par Météo-France. Il redistribue des données publiques sous Licence Ouverte 2.0. Source et attribution Source : Météo-France, données publiques de climatologie de base. Moissonnées depuis… See the full description on the dataset page: https://huggingface.co/datasets/cm3r/meteo-france-climatologie-quotidienne.tabular100M<n<1B0 likes270 downloads13d agoHugging Face18OpenLLM-France /Claire-Dialogue-English-0.1 Claire English Dialogue Dataset (CEDD) A collection of English dialogue transcripts This is the first packaged version of the datasets used to train the english variants of the Claire family of large language models (OpenLLM-France/Claire-7B-EN-0.1). (A related French dataset can be found here.) The Claire English Dialogue Dataset (CEDD) is a collection of transcripts of English dialogues from various sources, including parliamentary proceedings, interviews, broadcast, meetings, and… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Claire-Dialogue-English-0.1.texttext-generation100K<n<1M5 likes237 downloads2y agoHugging Face19francepfl /mantis_datasettext100K<n<1M0 likes232 downloads2y agoHugging Face20saget-antoine /francecrops FranceCrops FranceCrops is a crop-classification benchmark for French agricultural parcels observed with Sentinel-2 L2A time series meant to evaluate the representation learned by self-supervised or unsupervised methods. Each sample is one parcel represented by 100 sampled pixel time series. This release provides fixed supervised splits for downstream evaluation and frozen low-label subsets from 1 to 4,000 labels per class so methods can be compared under the same downstream… See the full description on the dataset page: https://huggingface.co/datasets/saget-antoine/francecrops.tabulartabular-classification100K<n<1M3 likes205 downloads3mo agoHugging Face21francescapistilli /sg-ego SG-Ego Annotations Project Page: https://francescapistilli.github.io/GLEN. Data annotation pipeline: https://github.com/francescapistilli/sg-ego SG-Ego is a large scale annotation set extending Ego4D with spatio-temporal scene graphs, where relations triplets are consolidated over time into explicit time-evolving descriptions of the scene state. SG-Ego is released as part of our paper "Learning to Evolve Scenes: Reasoning about Human Activities with Scene Graphs", which presents… See the full description on the dataset page: https://huggingface.co/datasets/francescapistilli/sg-ego.tabular10M<n<100M1 likes177 downloads3mo agoHugging Face22INPI-France /French-Patent-1981-2026-Clean 🇫🇷 Brevets français 1981–2026 — Clean 🇫🇷 Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet). Format : Parquet, prêt pour chargement streaming / distribué. Source Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI). Génération… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/French-Patent-1981-2026-Clean.tabular100K<n<1M1 likes172 downloads3mo agoHugging Face23INPI-France /Brevets-Francais-2024-Chunked 🇫🇷 Brevets français 2024 Chunké 🇫🇷 Dataset de brevets français publiés en 2024, extrait depuis les XML d’origine et chunké au niveau des balises <p> xml Format : Parquet, prêt pour chargement streaming / distribué. Source Données issues de documents publics de brevets français (A1, 2024).Extraction, structuration et découpage réalisés de manière indépendante grace a un acces aux API/FTP PI (sur demande a l'inpi). Génération Entrée : 35 479 fichiers XML… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/Brevets-Francais-2024-Chunked.tabular1M<n<10M1 likes155 downloads8mo agoHugging Face24meteolibre-dev /weather_france_radar_satellite_gs This dataset is a ready-to-use fusion of multiple datasets for the France region, including: Ground station data: Sourced from data.gouv.fr with a 1-hour resolution. It includes 7 weather KPIs: "RR1", "FF", "DD", "T", "U", "PMER", and "VV". Radar imagery: Radar rainfall accumulation images from the Météo-France open data initiative. Geospatial data: Land cover and ground height information from EarthEnv and OpenTopography. Satellite imagery: Sourced from the EUMETSAT platform… See the full description on the dataset page: https://huggingface.co/datasets/meteolibre-dev/weather_france_radar_satellite_gs.tabular10K<n<100K0 likes142 downloads1y agoHugging Face25OpenLLM-France /Claire-Dialogue-French-0.1gated Claire French Dialogue Dataset (CFDD) A collection of French dialogue transcripts and plays This is the first packaged version of the datasets used to train the Claire family of large language models (OpenLLM-France/Claire-7B-0.1). The Claire French Dialogue Dataset (CFDD) is a collection of theater plays and transcripts of real French dialogues from various sources, including parliamentary proceedings, interviews, debates, meetings, and free conversations. Each dialogue is split… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Claire-Dialogue-French-0.1.texttext-generation10K<n<100K57 likes130 downloads1y agoHugging Face26INPI-France /French-Patents-2020-2026-Raw 🇫🇷 Brevets français 2020–2026 — RAW 🇫🇷 Dataset de brevets français publiés entre 2020 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet). Format : Parquet, prêt pour chargement streaming / distribué. Source Données issues de documents publics de brevets français (A1).Extraction réalisé de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI). Génération 468 000 fichiers XML… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/French-Patents-2020-2026-Raw.text100K<n<1M3 likes126 downloads6mo agoHugging Face27tantara /Nemotron-Personas-France-Qwen3-0.6B-embedding Nemotron-Personas-France-Qwen3-0.6B-embedding Embeddings for nvidia/Nemotron-Personas-France computed with Qwen/Qwen3-Embedding-0.6B. Details Source dataset: nvidia/Nemotron-Personas-France Embedding model: Qwen/Qwen3-Embedding-0.6B Embedding dimension: 1024 Number of rows: 1000000 (first 1M rows of the 7M-row source dataset) Columns embedded (in dataset order): professional_persona, sports_persona, arts_persona, travel_persona, culinary_persona, persona… See the full description on the dataset page: https://huggingface.co/datasets/tantara/Nemotron-Personas-France-Qwen3-0.6B-embedding.text1M<n<10M0 likes114 downloads4mo agoHugging Face28francescortu /comp-mechEach record in the dataset contains the following fields: target_new: the counterfactual term target_true: the actual term subject: the topic of the prompt base_prompt: the foundational prompt prompt: the modified prompt incorporating the counterfactual change template: the sentence structure using the counterfactual wording. text10K<n<100K1 likes111 downloads2y agoHugging Face29justicedao /ipfs_france_laws_ir France legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_france_laws (revision ``) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of France prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was invented. Primary key: entry_cid (CIDv1 raw… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_france_laws_ir.tabulartext-retrieval10M<n<100M0 likes107 downloads1d agoHugging Face30OpenLLM-France /Luciole_RAG Dataset overview Luciole RAG is a supervised fine-tuning dataset for retrieval-augmented generation, built to train the Luciole models. Each example is a chat conversation where the assistant answers a question using only a set of retrieved document chunks given in the system prompt, quotes and cites its sources, and declines to answer when the documents do not contain the answer. It contains two subsets derived from existing question-answering benchmarks: Config Source… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole_RAG.textquestion-answering10K<n<100K1 likes104 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.