CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenLLM-France /Lucie-Training-Dataset Lucie Training Dataset Card The Lucie Training Dataset is a curated collection of text data in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers, digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages. The Lucie Training Dataset was used to pretrain Lucie-7B, a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.texttext-generation10B<n<100B39 likes30k downloads1y agoHugging Face02gollumeo /France_Government_Conversations0 likes9k downloads3y agoHugging Face03SauryChen /SeasonBench-EA-METEO_FRANCE-PressureThe dataset is a subset for SeasonBench-EA Benchmark, which contains the ensemble forecasts from Meteo France on pressure levels. All data are downloaded from the Copernicus Climate Data Store (https://cds.climate.copernicus.eu/) and reorganized for benchmark construction. Please ensure compliance with the CDS Licence when using or redistributing the data. 0 likes2.2k downloads1y agoHugging Face04OpenLLM-France /Luciole-Training-Dataset Data card for The Luciole Training Dataset Table of Contents Dataset Description Curation Rationale Web Data Opt-Outs Personal and Sensitive Information (PII) Bias, Risks, and Limitations Recommendations Sample Metadata Downloading the Data Sample Use in Python Accessing the English Web Data and OpenMathInstruct-1 Details on Data Sources Citation Acknowledgements Contact Dataset Description The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.texttext-generation1B<n<10B15 likes1.8k downloads2mo agoHugging Face05OpenLLM-France /Luciole-PostTraining-Dataset-1.1 Table of Contents Dataset Description Curation Rationale Bias, Risks, and Limitations Data Subsets Sample Metadata Downloading the Data Available Configurations Loading Examples Accessing Data Through the Directory Hierarchy Details on Data Sources Citation Acknowledgements Contact Dataset Description The Luciole-PostTraining-Dataset-1.1 is a curated collection of open, instruction-style text data designed for language model post training. It includes a mixture of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-PostTraining-Dataset-1.1.text1M<n<10M5 likes1.8k downloads16h agoHugging Face06INPI-France /Brevets-Francais-1981-2026-Clean 🇫🇷 Brevets français 1981–2026 — Clean 🇫🇷 Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet). Format : Parquet, prêt pour chargement streaming / distribué. Source Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI). Génération Entrée :… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/Brevets-Francais-1981-2026-Clean.tabular100K<n<1M1 likes1.1k downloads8mo agoHugging Face07OpenLLM-France /wikipedia Plain text of Wikipedia Dataset Description Size Example use (python) Data fields Notes on data formatting License Aknowledgements Citation Dataset Description This dataset is a plain text version of pages from wikipedia.org spaces for several languages (English, German, French, Spanish, Italian). The text is without HTML tags nor wiki templates. It just includes markdown syntax for headers, lists and tables. See Notes on data formatting for more details. It was… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikipedia.texttext-generation10M<n<100M5 likes1.1k downloads2y agoHugging Face08INPI-France /Brevets-Francais-2000-2026-Rawtext100K<n<1M1 likes1k downloads8mo agoHugging Face09INPI-France /Brevets-Francais-1981-2026-Raw 🇫🇷 Brevets français 1981–2026 — Raw 🇫🇷 Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet). Format : Parquet Source Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI). Génération Entrée : 822 310 fichiers XML Sortie : 822 310 lignes… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/Brevets-Francais-1981-2026-Raw.text100K<n<1M3 likes923 downloads5mo agoHugging Face10OpenLLM-France /wikimedia Dataset Card This dataset is a curated collection of Wikimedia pages in markdown format, compiled from various Wikimedia projects across multiple languages. Covered Wikimedia Projects: wikipedia wikibooks wikinews wikiquote wikisource wikiversity wikivoyage wiktionary Supported Languages: ar (Arabic) br (Breton) ca (Catalan) co (Corsican) de (German) en (English) es (Spanish) eu (Basque) fr (French) frp (Arpitan) it (Italian) nl (Dutch) oc (Occitan) pcd (Picard) pt (Portuguese)… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikimedia.texttext-generation10M<n<100M3 likes856 downloads1y agoHugging Face11nvidia /Nemotron-Personas-France Nemotron-Personas-France Une approche d'IA composée pour des personas ancrés dans des distributions réelles A compound AI approach to personas grounded in real-world distributions Vue d'ensemble du jeu de données (Dataset Overview) Nemotron-Personas-France est un jeu de données en libre accès (CC BY 4.0) composé de personas générés de manière synthétique. Ce jeu de données s'appuie sur les distributions démographiques, géographiques et de traits de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-France.imagetext-generation1M<n<10M89 likes837 downloads6mo agoHugging Face12annaelmoussa /horse-racing-france French Horse Racing Dataset Structured dataset of French horse racing covering 2014-01-01 → 2026-05-05, restricted to French meetings. This repository is provided "as is" under an other license. Overview Longitudinal collection of French race meetings with full race, runner, and payout records. The data is split into four normalized tables, each exposed as an independent Hub configuration. Configuration Description Rows reunions Race meetings (one row… See the full description on the dataset page: https://huggingface.co/datasets/annaelmoussa/horse-racing-france.tabularother1M<n<10M1 likes574 downloads2mo agoHugging Face13INPI-France /Brevets-Francais-2020-2026-Rawtext100K<n<1M1 likes528 downloads8mo agoHugging Face14Francel12 /nave-hidalgo-dataset Nave Hidalgo — dataset completo del vuelo de inspeccion Capturas originales, sin procesar, del vuelo de inspeccion de una nave industrial y su estacionamiento (27 jul 2026) con un DJI Matrice 4T. Fotografias RGB (*_V.JPG) 3,216 · 4032x3024 · ~10.9 GB Fotografias termicas (*_T.JPG) 3,215 · 1280x1024 · ~5.8 GB Coordenadas del sitio 20.054569, -99.311673 Emparejado RGB <-> termico Cada toma existe como par. Comparten el prefijo completo (marca… See the full description on the dataset page: https://huggingface.co/datasets/Francel12/nave-hidalgo-dataset.image1K<n<10K0 likes527 downloads2mo agoHugging Face15saadtaleb /france-climate-risk-drias-tracc-2023 France Climate Risk DRIAS TRACC-2023 Ce dataset regroupe des fichiers NetCDF issus de DRIAS TRACC-2023 pour la France métropolitaine, organisés pour un usage plus simple. Le dépôt contient des indicateurs climatiques : pour les niveaux de réchauffement +2.0 °C, +2.7 °C et +4.0 °C ; en produits multi-modèles ; en produits par modèle ; pour les valeurs absolues ; ainsi que pour la période de référence historique. Structure du dépôt Le dépôt est organisé par niveau de… See the full description on the dataset page: https://huggingface.co/datasets/saadtaleb/france-climate-risk-drias-tracc-2023.0 likes436 downloads6mo agoHugging Face16OpenLLM-France /RULER-luciole_tokenizer_128k-arab-regional_v2tabular10K<n<100K0 likes374 downloads10mo agoHugging Face17INPI-France /Brevets-Francais-2025-Claimstabular100K<n<1M1 likes373 downloads7mo agoHugging Face18SauryChen /SeasonBench-EA-METEO_FRANCE-SingleLevelThe dataset is a subset for SeasonBench-EA Benchmark, which contains the ensemble forecasts from Meteo France on single level. All data are downloaded from the Copernicus Climate Data Store (https://cds.climate.copernicus.eu/) and reorganized for benchmark construction. Please ensure compliance with the CDS Licence when using or redistributing the data. 0 likes360 downloads1y agoHugging Face19Francesco /who-killed-jfkimage10K<n<100K2 likes323 downloads1y agoHugging Face20OpenLLM-France /Translation-Instruct Corpus overview Translation Instruct is a collection of parallel corpora for machine translation, formatted as instructions for the supervised fine-tuning of large language models. It currently contains two collections: Croissant Aligned Instruct and Europarl Aligned Instruct. Croissant Aligned Instruct is an instruction-formatted version of the parallel French-English data in croissantllm/croissant_dataset_no_web_data (subset: aligned_36b). Europarl Aligned Instruct is an… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Translation-Instruct.texttext-generation100K<n<1M5 likes295 downloads1y agoHugging Face21OpenLLM-France /wikisource Plain text of Wikisource Dataset Description Size Example use (python) Data fields Notes on data formatting License Aknowledgements Citation Dataset Description This dataset is a plain text version of pages from wikisource.org in French language. The text is without HTML tags nor wiki templates. It just includes markdown syntax for headers, lists and tables. See Notes on data formatting for more details. It was created by LINAGORA and OpenLLM France from the… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikisource.texttext-generation100K<n<1M4 likes287 downloads2y agoHugging Face22cm3r /meteo-france-climatologie-quotidienne Climatologie quotidienne Météo-France Archive des observations climatologiques quotidiennes des stations Météo-France, de 1786-05-01 à 2026-09-10, en Parquet partitionné par département et par période. Redistribution non officielle. Ce dépôt n'est pas publié par Météo-France. Il redistribue des données publiques sous Licence Ouverte 2.0. Source et attribution Source : Météo-France, données publiques de climatologie de base. Moissonnées depuis… See the full description on the dataset page: https://huggingface.co/datasets/cm3r/meteo-france-climatologie-quotidienne.tabular100M<n<1B0 likes270 downloads13d agoHugging Face23francesco-immorlano /tl-climate-projections0 likes265 downloads3mo agoHugging Face24nicO1asFr /crosswalks-france-samimagen<1K0 likes260 downloads11mo agoHugging Face25Francesco /vehicles-q0x2v Dataset Card for vehicles-q0x2v ** The original COCO dataset is stored at dataset.tar.gz** Dataset Summary vehicles-q0x2v Supported Tasks and Leaderboards object-detection: The dataset can be used to train a model for Object Detection. Languages English Dataset Structure Data Instances A data point comprises an image and its object annotations. { 'image_id': 15, 'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/vehicles-q0x2v.imageobject-detection1K<n<10K8 likes255 downloads3y agoHugging Face26Francesco /axial-mri Dataset Card for axial-mri ** The original COCO dataset is stored at dataset.tar.gz** Dataset Summary axial-mri Supported Tasks and Leaderboards object-detection: The dataset can be used to train a model for Object Detection. Languages English Dataset Structure Data Instances A data point comprises an image and its object annotations. { 'image_id': 15, 'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=640x640… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/axial-mri.imageobject-detectionn<1K1 likes238 downloads3y agoHugging Face27OpenLLM-France /Claire-Dialogue-English-0.1 Claire English Dialogue Dataset (CEDD) A collection of English dialogue transcripts This is the first packaged version of the datasets used to train the english variants of the Claire family of large language models (OpenLLM-France/Claire-7B-EN-0.1). (A related French dataset can be found here.) The Claire English Dialogue Dataset (CEDD) is a collection of transcripts of English dialogues from various sources, including parliamentary proceedings, interviews, broadcast, meetings, and… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Claire-Dialogue-English-0.1.texttext-generation100K<n<1M5 likes237 downloads2y agoHugging Face28francepfl /mantis_datasettext100K<n<1M0 likes232 downloads2y agoHugging Face29nicO1asFr /pedestrian-crosswalks-france-v20 likes228 downloads11mo agoHugging Face30Francesco /insects-mytwu Dataset Card for insects-mytwu ** The original COCO dataset is stored at dataset.tar.gz** Dataset Summary insects-mytwu Supported Tasks and Leaderboards object-detection: The dataset can be used to train a model for Object Detection. Languages English Dataset Structure Data Instances A data point comprises an image and its object annotations. { 'image_id': 15, 'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/insects-mytwu.imageobject-detectionn<1K2 likes213 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.