CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HumynLabs /French_Documents_Dataset_PDF French Documents Dataset (PDF) This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.documentn<1K0 likes1.2k downloads11mo agoHugging Face02JDKdev /french-tts-conversational-dataset French Conversational TTS Dataset Dataset Description This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's Voxtral Mini TTS model (voxtral-mini-tts-2603). It covers three B2B industry verticals with balanced male/female speaker distribution. Verticals Vertical Description fintech_banking Banking operations, account inquiries, fraud alerts, investments, customer service ecommerce_logistics Order… See the full description on the dataset page: https://huggingface.co/datasets/JDKdev/french-tts-conversational-dataset.audiotext-to-speechn<1K0 likes916 downloads3mo agoHugging Face03AdrienB134 /Emilia-dataset-french-with-gender Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.audioautomatic-speech-recognition100K<n<1M1 likes639 downloads2y agoHugging Face04AdrienB134 /Emilia-dataset-french-splitaudio100K<n<1M4 likes421 downloads2y agoHugging Face05freds0 /cml_tts_dataset_frenchaudio100K<n<1M2 likes351 downloads2y agoHugging Face06french-datasets /sakthivinash-Language_DetectionCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données https://huggingface.co/datasets/sakthivinash/Language_Detection. 0 likes345 downloads1y agoHugging Face07Archime /french_tv_media_dataset_2026 Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus Résumé (Abstract) Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/Archime/french_tv_media_dataset_2026.audioautomatic-speech-recognition10K<n<100K4 likes343 downloads8mo agoHugging Face08madoss /french_tv_media_dataset_2026 Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus Résumé (Abstract) Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/madoss/french_tv_media_dataset_2026.audioautomatic-speech-recognition10K<n<100K1 likes270 downloads2mo agoHugging Face09manu /croissant_french_datasettext1M<n<10M0 likes264 downloads3y agoHugging Face10jpacifico /French-Alpaca-dataset-Instruct-110K110368 French instructions generated by OpenAI GPT-3.5-turbo in Alpaca Format to finetune general models Created by Jonathan Pacifico, 2024Please credit my name if you use this dataset in your project. text100K<n<1M14 likes166 downloads2y agoHugging Face11Tm24sense /english-french-datasettext10M<n<100M1 likes156 downloads1y agoHugging Face12jpacifico /French-Alpaca-dataset-Instruct-55K55184 french instructions generated by OpenAI GPT-3.5 in Alpaca Format to finetune general models Created by Jonathan Pacifico license: apache-2.0 Please credit my name if you use this dataset in your project. text10K<n<100K4 likes118 downloads3y agoHugging Face13MaroneAI /Wolof-to-French_Translation-Dataset Dataset Wolof ↔ Français 🧩 Présentation Ce dataset contient plus de 30 000 paires phrase Wolof – phrase Française.Chaque ligne est structurée comme suit : Wolof (input) Français (target) Phrase en Wolof Phrase correspondante en Français Il a été conçu pour la traduction automatique et les tâches de NLP impliquant le Wolof et le Français. 📚 Provenance et nettoyage Le dataset a été créé en compilant différentes sources accessibles… See the full description on the dataset page: https://huggingface.co/datasets/MaroneAI/Wolof-to-French_Translation-Dataset.texttranslation10K<n<100K3 likes100 downloads1y agoHugging Face14french-datasets /Makxxx-french_CEFRCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données huggingface.co/datasets/Makxxx/french_CEFR. 0 likes98 downloads1y agoHugging Face15french-datasets /vekkt-french_CEFRCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données huggingface.co/datasets/vekkt/french_CEFR. 0 likes93 downloads1y agoHugging Face16voxozi /french-tts-conversational-dataset French Conversational TTS Dataset Dataset Description This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's Voxtral Mini TTS model (voxtral-mini-tts-2603). It covers three B2B industry verticals with balanced male/female speaker distribution. Verticals Vertical Description fintech_banking Banking operations, account inquiries, fraud alerts, investments, customer service ecommerce_logistics Order… See the full description on the dataset page: https://huggingface.co/datasets/voxozi/french-tts-conversational-dataset.audiotext-to-speechn<1K0 likes71 downloads3mo agoHugging Face17rishabbahal /quebecois_canadian_french_datasetaudio1K<n<10K5 likes70 downloads2y agoHugging Face18vonewman /french-instruction-datasettext100K<n<1M1 likes68 downloads2y agoHugging Face19bilalfaye /english-wolof-french-datasettext10K<n<100K0 likes63 downloads2y agoHugging Face20shunyalabs /french-speech-datasetaudio1K<n<10K0 likes58 downloads1y agoHugging Face21PeggyVallin /jfv-french-style-conditioning-dataset-v1.0 JFV French Paired Style-Conditioning Dataset At a Glance Item Value Language French Source Single-author blog corpus, 2005–2025 Public release v1.0 Aligned units in public release 1,484 Texts in public aligned release 7,420 Original experiment 1,492 aligned units / 7,460 texts Generated conditions Ministral baseline; profile; profile + five-shot examples Primary use Paired study of stylistic conditioning and evaluation-metric validity… See the full description on the dataset page: https://huggingface.co/datasets/PeggyVallin/jfv-french-style-conditioning-dataset-v1.0.tabulartext-generation1K<n<10K0 likes54 downloads5d agoHugging Face22AxonData /french-call-center-speech-dataset French Call Center Speech Dataset: 1,000+ Hours with Transcripts 1,000+ hours of real-world French call center audio with transcripts. Train speech recognition, sentiment analysis, and customer support AI models on authentic telephone conversations Dataset Summary Key Features ✅ 1,000+ hours of inbound & outbound calls✅ 100% French telephone conversations✅ Real-world audio - no synthetic data✅ Full transcripts in French and in English Full… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/french-call-center-speech-dataset.audion<1K0 likes47 downloads7mo agoHugging Face23CATIE-AQ /facebook-community-alignment-dataset_french_dpo Description This is the Community Alignment dataset which we've cleaned up to keep only the French datas (+ deduplication) and reformatted for DPO finetuning.For more details on the dataset itself, please consult the original dataset card or the paper. Original authors @article{zhang2025cultivating, title = {Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset}, author = {Lily Hong Zhang and Smitha Milli and Karen Jusko and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/facebook-community-alignment-dataset_french_dpo.text10K<n<100K1 likes43 downloads1y agoHugging Face24InfoBayAI /French_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 31,106 hours of processed French (FR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format, where… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/French_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes40 downloads9d agoHugging Face25CATIE-AQ /facebook-community-alignment-dataset_french_conversation Description This is the Community Alignment dataset which we've cleaned up to keep only the French datas (+ deduplication) and reformatted as a conversation to simplify his use for alignment finetuning.For more details on the dataset itself, please consult the original dataset card or the paper. Original authors @article{zhang2025cultivating, title = {Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset}, author = {Lily Hong Zhang… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/facebook-community-alignment-dataset_french_conversation.text10K<n<100K1 likes39 downloads1y agoHugging Face26Mozilla /query-intent-detection-dataset-frenchtext100K<n<1M0 likes38 downloads2mo agoHugging Face27davidpistori /mistral-legal-french-dataset Mistral Legal French Dataset A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy. 📋 Table of Contents Overview Dataset Composition Methodology 1. Chain-of-Thought Generation 2. LegalKit Extraction 3. Curriculum Learning Fusion Data Format Quality Metrics Usage Citations License 🎯 Overview This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two… See the full description on the dataset page: https://huggingface.co/datasets/davidpistori/mistral-legal-french-dataset.texttext-generation10K<n<100K0 likes37 downloads3mo agoHugging Face28Speech-data /French-Speech-Dataset 🎧 French Speech Dataset The French Speech Dataset is a comprehensive speech audio dataset designed to deliver high-quality and diverse audio data for advanced AI and machine learning applications. It includes 198 hours of audio data across 912 files, provided in MP3 and WAV formats, with a total size of 445 MB. This well-structured audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a wide age distribution from 18 to 50+ years.… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/French-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes36 downloads6mo agoHugging Face29Alijeff1214 /DILA_FRENCH_DATASETtext0 likes34 downloads2y agoHugging Face30Volko76 /smol-smoltalk-french-instruction-dataset French Instruction Dataset for Nanochat This dataset is a transformed version of vonewman/french-instruction-dataset. It has been formatted to match the SmolTalk structure required by nanochat. Format Format: ChatML / SmolTalk Column: messages (List of dicts with role and content) Usage with Nanochat self.ds = load_dataset("Volko76/smol-smoltalk-french-instruction-dataset", split=split) text100K<n<1M1 likes34 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.