CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PleIAs /French-PD-Newspapers 🇫🇷 French Public Domain Newspapers 🇫🇷 French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.tabulartext-generation1M<n<10M70 likes3.5k downloads3y agoHugging Face02PleIAs /French-PD-Books 🇫🇷 French Public Domain Books 🇫🇷 French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.tabulartext-generation100K<n<1M52 likes3k downloads3y agoHugging Face03PleIAs /French-Science-Commons French Science Commons French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres. Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.tabular10M<n<100M26 likes1.2k downloads3mo agoHugging Face04PleIAs /French-PD-diverse43,085,129,931 words tabular100K<n<1M3 likes689 downloads2y agoHugging Face05PhysiQuanty /FRENCH-ONLY-Common-Crawl-2026-25tabular1M<n<10M3 likes625 downloads3mo agoHugging Face06sprinklr-huggingface /CXM_Arena_French Dataset Card for CXM Arena French Benchmark Suite Dataset Description This dataset, "CXM Arena French Benchmark Suite," is a comprehensive collection designed to evaluate various AI capabilities within the Customer Experience Management (CXM) domain, specifically for the French language. It is closely modeled after the original CXM_Arena benchmark, but all data is in French. The suite consolidates five distinct tasks into a unified benchmark, enabling robust testing of… See the full description on the dataset page: https://huggingface.co/datasets/sprinklr-huggingface/CXM_Arena_French.document10K<n<100K1 likes513 downloads1y agoHugging Face07hamouda /French-PD-diverse43,085,129,931 words tabular100K<n<1M0 likes415 downloads7mo agoHugging Face08AIML-TUDA /SLR-Bench-French 🧠 SLR-Bench-French: Scalable Logical Reasoning Benchmark (French Edition) SLR-Bench Multilingual Versions: SLR-Bench-French is the French-language pendant of the original SLR-Benchdataset. It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into French. This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-French.tabular10K<n<100K1 likes387 downloads4mo agoHugging Face09harvard-lil /cold-french-law Collaborative Open Legal Data (COLD) - French Law COLD French Law is a dataset containing over 800 000 french law articles, filtered and extracted from France's LEGI dataset and formatted as a single CSV file. This dataset focuses on articles (codes, lois, décrets, arrêtés ...) identified as currently applicable french law. A large portion of this dataset comes with machine-generated english translations, provided by Casetext, Part of Thomson Reuters using OpenAI's GPT-4. This… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-french-law.tabular100K<n<1M21 likes340 downloads2y agoHugging Face10Abirate /french_book_reviews Dataset Card for French book reviews I-Dataset Summary The majority of review datasets are in English. There are datasets in other languages, but not many. Through this work, I would like to enrich the datasets in the French language(my mother tongue with Arabic).The data was retrieved from two French websites: Babelio and Critiques LibresLike Wikipedia, these two French sites are made possible by the contributions of volunteers who use the Internet to share their… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/french_book_reviews.tabulartext-classification1K<n<10K8 likes275 downloads4y agoHugging Face11bowang0911 /paraphrasing-french Attribution MTEB-format derivative of ismailiismail/paraphrasing_french. Query = phrase; corpus = paraphrase. tabulartext-retrieval1K<n<10K0 likes244 downloads3mo agoHugging Face12bowang0911 /alpaca-french-mixtral License & Attribution MTEB-format derivative of AIffl/Alpaca_french_mixtral (French Alpaca, Mixtral-translated). Query = instruction; corpus = answer. Deterministically subsampled to ~10k. Licensed under Apache-2.0 (same as source). tabulartext-retrieval10K<n<100K0 likes229 downloads3mo agoHugging Face13FrenchCastle /IMF-Reports IMF Technical Assistance Reports — Recommendation Process Corpus A page-grounded research corpus of 780 IMF technical-assistance report records. It contains source PDFs, layout-aware Markdown, page-level text, extracted visuals, metadata, observations, recommendations, and labeled links between observations and recommendations. Required acknowledgement All research, publications, datasets, models, applications, or other work derived from this corpus should… See the full description on the dataset page: https://huggingface.co/datasets/FrenchCastle/IMF-Reports.imagetext-classification100K<n<1M0 likes221 downloads2mo agoHugging Face14artefactory /Argimi-Legal-French-Jurisprudence The ArGiMi French Jurisprudence Dataset This dataset contains a comprehensive collection of French case law, sourced from the official archives of French jurisprudence. It is divided into three distinct subdivisions: Constitutional ("constit"), Administrative ("cetat"), and Judiciary ("juri"). This dataset was created for the ArGiMi project, an open-source initiative dedicated to promoting open data and knowledge sharing. The project is a collaborative effort between Giskard… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Legal-French-Jurisprudence.tabularquestion-answering100K<n<1M10 likes217 downloads1y agoHugging Face15INPI-France /French-Patent-1981-2026-Clean 🇫🇷 Brevets français 1981–2026 — Clean 🇫🇷 Dataset de brevets français publiés entre 1981 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet). Format : Parquet, prêt pour chargement streaming / distribué. Source Données issues de documents publics de brevets français (A1).Extraction, structuration et nettoyage réalisés de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI). Génération… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/French-Patent-1981-2026-Clean.tabular100K<n<1M1 likes186 downloads3mo agoHugging Face16krishnakamath /fama_french_datatabular1K<n<10K0 likes185 downloads9mo agoHugging Face17tadad /french-fiction-16-18th-century French Fiction of the 16th–18th Centuries A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model. The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction. Structure Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.tabulartext-classification100K<n<1M0 likes163 downloads18d agoHugging Face18Kartmaan /french-dictionary French Dictionary A ready-to-use offline French language dictionary derived from the French Wiktionary. Available in two formats to suit different use cases: SQLite for desktop applications and real-time querying, and Parquet for data science and machine learning pipelines. Contains nearly 900,000 distinct word forms including conjugated verb forms, with structured definitions, usage examples, and rich linguistic metadata. Acknowledgements This dataset would not… See the full description on the dataset page: https://huggingface.co/datasets/Kartmaan/french-dictionary.tabular1M<n<10M1 likes149 downloads6mo agoHugging Face19Jaymerry /french-administrative-hierarchy-data-quality French Administrative Hierarchy Data Quality Benchmark A reproducible benchmark for evaluating the validation, classification, and repair of French administrative geographic records. The dataset is derived from the INSEE Code officiel géographique (COG) 2026 and focuses on the hierarchical relationship: Region → Department → Commune Important: commune_code is an INSEE/COG administrative identifier, not a postal code.This dataset is not intended for postal-address validation or… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/french-administrative-hierarchy-data-quality.tabulartabular-classification100K<n<1M0 likes124 downloads2mo agoHugging Face20FrenchCastle /isora-tax-administration ISORA — International Survey on Revenue Administration, FY2014–FY2024 Every published answer of every ISORA survey round, in one clean long-format panel, with the metadata you need to use it responsibly: what each question means in each questionnaire generation, which questions changed wording (or meaning) between rounds, which jurisdictions answered which question in which year, and how published values were revised between releases. ISORA is the joint survey of national tax… See the full description on the dataset page: https://huggingface.co/datasets/FrenchCastle/isora-tax-administration.tabulartabular-classification100K<n<1M0 likes110 downloads1d agoHugging Face21maximoss /rte3-french Dataset Card for French RTE-3 Dataset Summary The RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge. Like its English counterpart, the French RTE-3 dataset is composed of a development set and a test set, each containing 800 T/H pairs. All T/H pairs were manually translated into French and proofread. It is annotated for a 3-way task. Please refer to this repository, if you want to use this French version… See the full description on the dataset page: https://huggingface.co/datasets/maximoss/rte3-french.tabulartext-classification1K<n<10K0 likes104 downloads7mo agoHugging Face22Paul /hatecheck-french Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-french.tabulartext-classification1K<n<10K0 likes101 downloads4y agoHugging Face23Nicolas-BZRD /English_French_Songs_Lyrics_Translation_Original Original Songs Lyrics with French Translation Dataset Summary Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French. Details of the number of songs by language of origin can be found in the table below: Original language Number of songs en 75786 fr 18486 es 1743 it 803 de 691 sw 529 ko 193 id 169 pt 142 no 122 fi 113 sv 70 hr 53 so 43 ca 41 tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.tabulartranslation10K<n<100K16 likes95 downloads3y agoHugging Face24rntc /clinical-qa-french-v2tabular100K<n<1M1 likes88 downloads8mo agoHugging Face25VALORISIMMO /valoris-french-real-estate-prices French Real Estate Prices — VALORIS Observatory Aggregated real estate median prices (€/m²) computed from the French open data source DVF (Demandes de Valeurs Foncières) published by DGFiP. Covers 93 departments and their communes of metropolitan France (excluding Alsace-Moselle departments 57, 67, 68 — local Livre Foncier system). 🔗 Interactive visualization & drill-down: valoris-immo.fr/observatoire 🏠 Publisher homepage: valoris-immo.fr 📊 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/VALORISIMMO/valoris-french-real-estate-prices.tabulartabular-regression100K<n<1M0 likes82 downloads5mo agoHugging Face26mfgiguere /erudit-french-philosophy Dataset Card for Dataset Name Dataset Description Dataset Summary This dataset contains all french philosophy that has been published on erudit.org. It has been generated using a Bs4 web parser that you can find in this repo: https://github.com/MFGiguere/french-philosophy-generator. Supported Tasks and Leaderboards This dataset could be useful for this (non-exhaustive) set of tasks: detect if a text is philosophical or not, generate philosophical… See the full description on the dataset page: https://huggingface.co/datasets/mfgiguere/erudit-french-philosophy.tabular100K<n<1M2 likes69 downloads3y agoHugging Face27frenchtext /banque-fr-2311 Dataset Card for "banque fr websites - 2311" Dataset extracted from public websites by wordslab-webscraper in 2311: domain: banque language: fr license: Apache 2.0 Dataset Sources wordslab-webscraper follows the industry best practices for polite web scraping: clearly identifies itself as a known text indexing bot: "bingbot" doesn't try to hide the user IP address behind proxies doesn't try to circumvent bots protection solutions waits for a minimum delay between two… See the full description on the dataset page: https://huggingface.co/datasets/frenchtext/banque-fr-2311.tabulartext-generation10K<n<100K0 likes69 downloads3y agoHugging Face28ahmad21omar /SLR-Bench-French 🧠 SLR-Bench-French: Scalable Logical Reasoning Benchmark (French Edition) SLR-Bench Versions: SLR-Bench-French is the French-language pendant of the original SLR-Bench dataset. It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into French. This enables systematic evaluation and training of Large Language Models (LLMs) in logical reasoning in French… See the full description on the dataset page: https://huggingface.co/datasets/ahmad21omar/SLR-Bench-French.tabular10K<n<100K0 likes67 downloads11mo agoHugging Face29louisbertson /french-moore-parallel French → Mooré (Mossi) Parallel Corpus Machine-translated parallel sentences from French (fr) to Mooré / Mossi (mos), produced by a public-web crawl + filtering + Glosbe translation pipeline. Snapshot Field Value Validated pairs 3,000,040 Source language French Target language Mooré (Mossi) Translator Glosbe public MT Export date 2026-08-14 Schema Column Type Description id string (UUID) Pair identifier… See the full description on the dataset page: https://huggingface.co/datasets/louisbertson/french-moore-parallel.tabulartranslation1M<n<10M1 likes67 downloads1mo agoHugging Face30FrancophonIA /french_financial_news [!NOTE] Dataset origin: https://www.kaggle.com/datasets/arcticgiant/french-financial-news Context This dataset contains around 41 500 french news from 11/2018 to 03/2021 scraped on a famous financial media website. For ease of use I’v add English translation (Helsinki-NLP/opus-mt-fr-en) and sentiment analysis (VADER) Analysis The picture below show the effect of covid crisis on news sentiment (Purple) and CAC40 (Blue). We see clearly a link between the news sentiment… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/french_financial_news.tabular10K<n<100K1 likes59 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.