CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PleIAs /French-PD-Books 🇫🇷 French Public Domain Books 🇫🇷 French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.tabulartext-generation100K<n<1M52 likes3k downloads3y agoHugging Face02PleIAs /French-PD-Newspapers 🇫🇷 French Public Domain Newspapers 🇫🇷 French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.tabulartext-generation1M<n<10M70 likes2.7k downloads3y agoHugging Face03vekkt /french_CEFRtext1K<n<10K3 likes2.3k downloads4y agoHugging Face04britllm /TransWeb-Edu-Frenchtext10M<n<100M0 likes1.4k downloads2y agoHugging Face05PleIAs /French-Science-Commons French Science Commons French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres. Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.tabular10M<n<100M26 likes1.2k downloads3mo agoHugging Face06HumynLabs /French_Documents_Dataset_PDF French Documents Dataset (PDF) This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.documentn<1K0 likes1.2k downloads11mo agoHugging Face07the-french-artist /hatvp_raw_archivestextn<1K0 likes1k downloads2y agoHugging Face08Makxxx /french_CEFRtext1K<n<10K0 likes988 downloads4y agoHugging Face09manu /french-30b Dataset Card for "french_30b2" More Information needed text10M<n<100M0 likes965 downloads3y agoHugging Face10AdoCleanCode /SPEEED_s3_words_french_100k-300ktext100K<n<1M0 likes938 downloads8mo agoHugging Face11JDKdev /french-tts-conversational-dataset French Conversational TTS Dataset Dataset Description This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's Voxtral Mini TTS model (voxtral-mini-tts-2603). It covers three B2B industry verticals with balanced male/female speaker distribution. Verticals Vertical Description fintech_banking Banking operations, account inquiries, fraud alerts, investments, customer service ecommerce_logistics Order… See the full description on the dataset page: https://huggingface.co/datasets/JDKdev/french-tts-conversational-dataset.audiotext-to-speechn<1K0 likes907 downloads3mo agoHugging Face12Snit /french-conversation+15 hours of speech data from TTS and text file recording. +9k utterances from various sources, novels, parliamentary debates, professional language. audio10K<n<100K6 likes822 downloads3y agoHugging Face13bofenghuang /mt-bench-french MT-Bench-French This is a French version of MT-Bench, created to evaluate the multi-turn conversation and instruction-following capabilities of LLMs. Similar to its original version, MT-Bench-French comprises 80 high-quality, multi-turn questions spanning eight main categories. All questions have undergone translation into French and thorough human review to guarantee the use of suitable and authentic wording, meaningful content for assessing LLMs' capabilities in the French… See the full description on the dataset page: https://huggingface.co/datasets/bofenghuang/mt-bench-french.textquestion-answeringn<1K8 likes731 downloads2y agoHugging Face14PleIAs /French-PD-diverse43,085,129,931 words tabular100K<n<1M3 likes714 downloads2y agoHugging Face15Rcarvalo /french-dialogue-tts-1000haudio10K<n<100K1 likes674 downloads6mo agoHugging Face16PhysiQuanty /FRENCH-ONLY-Common-Crawl-2026-25tabular1M<n<10M3 likes598 downloads3mo agoHugging Face17AdoCleanCode /SPEEED_s3_words_french_0k-100k Multilingual Audio Alignments - Processed (Mixed Text/Phonemes) This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (french). Curriculum Learning This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule: p_start: 0.0 (starting probability of using phonemes) p_end: 0.0 (ending probability of using phonemes) curriculum_rows: 400000 (rows over which probability increases) Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_french_0k-100k.text100K<n<1M0 likes529 downloads8mo agoHugging Face18sprinklr-huggingface /CXM_Arena_French Dataset Card for CXM Arena French Benchmark Suite Dataset Description This dataset, "CXM Arena French Benchmark Suite," is a comprehensive collection designed to evaluate various AI capabilities within the Customer Experience Management (CXM) domain, specifically for the French language. It is closely modeled after the original CXM_Arena benchmark, but all data is in French. The suite consolidates five distinct tasks into a unified benchmark, enabling robust testing of… See the full description on the dataset page: https://huggingface.co/datasets/sprinklr-huggingface/CXM_Arena_French.document10K<n<100K1 likes514 downloads1y agoHugging Face19angeluriot /french_instruct 🧑‍🏫 French Instruct The French Instruct dataset is a collection of instructions with their corresponding answers (sometimes multi-turn conversations) entirely in French. The dataset is also available on GitHub. 📊 Overview The dataset is composed of 276K conversations between a user and an assistant for a total of approximately 85M tokens. I also added annotations for each document to indicate if it was generated or written by a human, the style of… See the full description on the dataset page: https://huggingface.co/datasets/angeluriot/french_instruct.textquestion-answering100K<n<1M19 likes480 downloads2y agoHugging Face20lbourdois /caption-wit_base_french Description This dataset is a processed version of wikimedia/wit_base.We converted the images to PIL format and kept only the lines containing French.The French texts come either from the caption_attribution_description column of the original dataset which we have renamed caption here, or from the wit_features column which we have renamed descriptions.The descriptions column is a list because it can contain several texts itself. For example, line 8 of the dataset contains three… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/caption-wit_base_french.imagevisual-question-answering1M<n<10M1 likes438 downloads8mo agoHugging Face21hamouda /French-PD-diverse43,085,129,931 words tabular100K<n<1M0 likes438 downloads7mo agoHugging Face22AdrienB134 /Emilia-dataset-french-splitaudio100K<n<1M4 likes422 downloads2y agoHugging Face23FreedomIntelligence /MMLU_FrenchFrench version of MMLU dataset tranlasted by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT. 2 likes421 downloads3y agoHugging Face24manu /french_5p Dataset Card for "french_5p" More Information needed text10M<n<100M0 likes403 downloads3y agoHugging Face25AdrienB134 /Emilia-dataset-french-with-gender Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.audioautomatic-speech-recognition100K<n<1M1 likes393 downloads2y agoHugging Face26AIML-TUDA /SLR-Bench-French 🧠 SLR-Bench-French: Scalable Logical Reasoning Benchmark (French Edition) SLR-Bench Multilingual Versions: SLR-Bench-French is the French-language pendant of the original SLR-Benchdataset. It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into French. This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-French.tabular10K<n<100K1 likes390 downloads4mo agoHugging Face27aai530-group6 /ddxplus-french Dataset Description We are releasing under the CC-BY licence a new large-scale dataset for Automatic Symptom Detection (ASD) and Automatic Diagnosis (AD) systems in the medical domain. The dataset contains patients synthesized using a proprietary medical knowledge base and a commercial rule-based AD system. Patients in the dataset are characterized by their socio-demographic data, a pathology they are suffering from, a set of symptoms and antecedents related to this pathology, and a… See the full description on the dataset page: https://huggingface.co/datasets/aai530-group6/ddxplus-french.texttabular-classification1M<n<10M2 likes387 downloads3y agoHugging Face28freds0 /cml_tts_dataset_frenchaudio100K<n<1M2 likes381 downloads2y agoHugging Face29roettger /eighteenth_century_french_novels General information This dataset contains 12 Mio Token of Literary French prose 1751-1800 in plain text format, built within the project 'Mining and Modeling Text' (2019-2023) at Trier University. For the dataset in XML/TEI see the GitHub repository of the project. Collection de romans français du dix-huitième siècle (1751-1800) / Collection of Eighteenth-Century French Novels (1751-1800) This collection of Eighteenth-Century French Novels contains 200 digital… See the full description on the dataset page: https://huggingface.co/datasets/roettger/eighteenth_century_french_novels.texttext-generation100K<n<1M2 likes352 downloads2y agoHugging Face30Rcarvalo /french-dialogue-tts-100haudio0 likes342 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.