CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PleIAs /French-PD-Newspapers 🇫🇷 French Public Domain Newspapers 🇫🇷 French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.tabulartext-generation1M<n<10M70 likes5.6k downloads3y agoHugging Face02PleIAs /French-PD-Books 🇫🇷 French Public Domain Books 🇫🇷 French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.tabulartext-generation100K<n<1M52 likes3k downloads3y agoHugging Face03angeluriot /french_instruct 🧑‍🏫 French Instruct The French Instruct dataset is a collection of instructions with their corresponding answers (sometimes multi-turn conversations) entirely in French. The dataset is also available on GitHub. 📊 Overview The dataset is composed of 276K conversations between a user and an assistant for a total of approximately 85M tokens. I also added annotations for each document to indicate if it was generated or written by a human, the style of… See the full description on the dataset page: https://huggingface.co/datasets/angeluriot/french_instruct.textquestion-answering100K<n<1M19 likes472 downloads2y agoHugging Face04roettger /eighteenth_century_french_novels General information This dataset contains 12 Mio Token of Literary French prose 1751-1800 in plain text format, built within the project 'Mining and Modeling Text' (2019-2023) at Trier University. For the dataset in XML/TEI see the GitHub repository of the project. Collection de romans français du dix-huitième siècle (1751-1800) / Collection of Eighteenth-Century French Novels (1751-1800) This collection of Eighteenth-Century French Novels contains 200 digital… See the full description on the dataset page: https://huggingface.co/datasets/roettger/eighteenth_century_french_novels.texttext-generation100K<n<1M2 likes359 downloads2y agoHugging Face05artefactory /Argimi-Legal-French-Jurisprudence The ArGiMi French Jurisprudence Dataset This dataset contains a comprehensive collection of French case law, sourced from the official archives of French jurisprudence. It is divided into three distinct subdivisions: Constitutional ("constit"), Administrative ("cetat"), and Judiciary ("juri"). This dataset was created for the ArGiMi project, an open-source initiative dedicated to promoting open data and knowledge sharing. The project is a collaborative effort between Giskard… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Legal-French-Jurisprudence.tabularquestion-answering100K<n<1M10 likes192 downloads1y agoHugging Face06Kant1 /French_Wikipedia_articlesDump of 2023-08-20 of all french article in wikipedia https://dumps.wikimedia.org/frwiki/20230820/frwiki-20230820-pages-articles.xml.bz2 texttext-generation10M<n<100M3 likes184 downloads3y agoHugging Face07tadad /french-fiction-16-18th-century French Fiction of the 16th–18th Centuries A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model. The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction. Structure Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.tabulartext-classification100K<n<1M0 likes166 downloads19d agoHugging Face08OpenLLM-France /Claire-Dialogue-French-0.1gated Claire French Dialogue Dataset (CFDD) A collection of French dialogue transcripts and plays This is the first packaged version of the datasets used to train the Claire family of large language models (OpenLLM-France/Claire-7B-0.1). The Claire French Dialogue Dataset (CFDD) is a collection of theater plays and transcripts of real French dialogues from various sources, including parliamentary proceedings, interviews, debates, meetings, and free conversations. Each dialogue is split… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Claire-Dialogue-French-0.1.texttext-generation10K<n<100K57 likes128 downloads1y agoHugging Face09Nicolas-BZRD /English_French_Songs_Lyrics_Translation_Original Original Songs Lyrics with French Translation Dataset Summary Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French. Details of the number of songs by language of origin can be found in the table below: Original language Number of songs en 75786 fr 18486 es 1743 it 803 de 691 sw 529 ko 193 id 169 pt 142 no 122 fi 113 sv 70 hr 53 so 43 ca 41 tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.tabulartranslation10K<n<100K16 likes99 downloads3y agoHugging Face10Lots-of-LoRAs /task815_pawsx_japanese_french_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task815_pawsx_japanese_french_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task815_pawsx_japanese_french_translation.texttext-generationn<1K0 likes78 downloads2y agoHugging Face11Archangel-system /french-dev-questions-5k French Developer Questions 5K 5,000 software-engineering questions written in French, generated by 22 open models across a combinatorial seed grid, then filtered one by one by an LLM judge. No answers — this is a prompt corpus, meant to be the input side of a distillation or SFT pipeline. French technical data is scarce. Most French datasets on the Hub are literary, legal or journalistic corpora; most developer-question datasets are English-only. This one sits in the… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/french-dev-questions-5k.texttext-generation1K<n<10K1 likes75 downloads9d agoHugging Face12frenchtext /banque-fr-2311 Dataset Card for "banque fr websites - 2311" Dataset extracted from public websites by wordslab-webscraper in 2311: domain: banque language: fr license: Apache 2.0 Dataset Sources wordslab-webscraper follows the industry best practices for polite web scraping: clearly identifies itself as a known text indexing bot: "bingbot" doesn't try to hide the user IP address behind proxies doesn't try to circumvent bots protection solutions waits for a minimum delay between two… See the full description on the dataset page: https://huggingface.co/datasets/frenchtext/banque-fr-2311.tabulartext-generation10K<n<100K0 likes70 downloads3y agoHugging Face13Lots-of-LoRAs /task778_pawsx_english_french_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task778_pawsx_english_french_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task778_pawsx_english_french_translation.texttext-generationn<1K0 likes64 downloads2y agoHugging Face14AIffl /french_orca_dpo_pairs Dataset Card for french_orca_dpo_pairs This dataset offers a french translation of the 12k DPO Intel/orca_dpo_pairs pairs made from Open-Orca/OpenOrca. Dataset Card Contact ntnq texttext-generation10K<n<100K7 likes62 downloads2y agoHugging Face15frenchtext /bank-en-2401 Dataset Card for "bank en websites - 2401" Dataset extracted from public websites by wordslab-webscraper in 2401: domain: bank language: en license: Apache 2.0 Dataset Sources wordslab-webscraper follows the industry best practices for polite web scraping: clearly identifies itself as a known text indexing bot: "bingbot" doesn't try to hide the user IP address behind proxies doesn't try to circumvent bots protection solutions waits for a minimum delay between two pages… See the full description on the dataset page: https://huggingface.co/datasets/frenchtext/bank-en-2401.tabulartext-generation10K<n<100K0 likes54 downloads3y agoHugging Face16PeggyVallin /jfv-french-style-conditioning-dataset-v1.0 JFV French Paired Style-Conditioning Dataset At a Glance Item Value Language French Source Single-author blog corpus, 2005–2025 Public release v1.0 Aligned units in public release 1,484 Texts in public aligned release 7,420 Original experiment 1,492 aligned units / 7,460 texts Generated conditions Ministral baseline; profile; profile + five-shot examples Primary use Paired study of stylistic conditioning and evaluation-metric validity… See the full description on the dataset page: https://huggingface.co/datasets/PeggyVallin/jfv-french-style-conditioning-dataset-v1.0.tabulartext-generation1K<n<10K0 likes54 downloads4d agoHugging Face17CATIE-AQ /frenchPARAPHRASE Dataset information Dataset concatenating Paraphases datasets available in French and open-source.There are a total of 254,513 rows, of which 251,753 are for training, 1,857 for validation and 903 for testing. Usage from datasets import load_dataset dataset = load_dataset("CATIE-AQ/frenchPARAPHRASE") Dataset Details of rows Dataset Original Splits Note Helsinki-NLP/tatoeba_mt 2,117 train / 999 validation We only keep the French split… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/frenchPARAPHRASE.texttext-generation100K<n<1M1 likes41 downloads10mo agoHugging Face18CreativeAlloyYT /French_Grammar_Explanations This dataset contains 1500+ French grammar explanations. It's the one I used to train my finetuned LLM called FrenchLlama-3.2-1B-Instruct. You can use this dataset for your own training purposes & find the aforementioned model on my HuggingFace profile. textquestion-answering1K<n<10K0 likes38 downloads1y agoHugging Face19TheJeanneCompany /french-senate-session-reports 🏛️ French Senate Session Reports Dataset A dataset of parliamentary debates and sessions reports from the French Senate.508,647,861 tokens of high-quality French text transcribed manually from Senate Sessions Description This dataset consists of all session reports from the French Senate debates, crawled from the official website senat.fr. It provides high-quality text data of parliamentary discussions, covering a wide range of political, economic, and social topics… See the full description on the dataset page: https://huggingface.co/datasets/TheJeanneCompany/french-senate-session-reports.texttext-classification1K<n<10K4 likes38 downloads2y agoHugging Face20Lots-of-LoRAs /task483_cls_french_dvd_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task483_cls_french_dvd_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task483_cls_french_dvd_classification.texttext-generation1K<n<10K0 likes37 downloads2y agoHugging Face21davidpistori /mistral-legal-french-dataset Mistral Legal French Dataset A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy. 📋 Table of Contents Overview Dataset Composition Methodology 1. Chain-of-Thought Generation 2. LegalKit Extraction 3. Curriculum Learning Fusion Data Format Quality Metrics Usage Citations License 🎯 Overview This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two… See the full description on the dataset page: https://huggingface.co/datasets/davidpistori/mistral-legal-french-dataset.texttext-generation10K<n<100K0 likes37 downloads3mo agoHugging Face22frenchtext /bank-de-2401 Dataset Card for "bank de websites - 2401" Dataset extracted from public websites by wordslab-webscraper in 2401: domain: bank language: de license: Apache 2.0 Dataset Sources wordslab-webscraper follows the industry best practices for polite web scraping: clearly identifies itself as a known text indexing bot: "bingbot" doesn't try to hide the user IP address behind proxies doesn't try to circumvent bots protection solutions waits for a minimum delay between two pages… See the full description on the dataset page: https://huggingface.co/datasets/frenchtext/bank-de-2401.tabulartext-generation10K<n<100K0 likes36 downloads3y agoHugging Face23AIffl /Alpaca_french_mixtral Dataset Card for Alpaca_french_mixtral This dataset was made by reusing the french alpaca instruction with Mixtral-8x7B-Instruct to make the output open-source. Dataset Card Contact robinjo textquestion-answering10K<n<100K4 likes34 downloads2y agoHugging Face24CGCTG /french_qa Wiki-FR-QA : Dataset de Questions-Réponses en Français Description Dataset de question-answering en français généré automatiquement à partir d'articles Wikipedia FR. Les questions sont produites par Qwen3.5-4B et filtrées pour la qualité. Structure Champ Description id Identifiant unique (SHA-256 tronqué) context Section Wikipedia servant de contexte question Question en français answer Réponse en une phrase article_title Article source… See the full description on the dataset page: https://huggingface.co/datasets/CGCTG/french_qa.textquestion-answering10K<n<100K1 likes34 downloads7mo agoHugging Face25Lots-of-LoRAs /task484_cls_french_music_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task484_cls_french_music_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task484_cls_french_music_classification.texttext-generation1K<n<10K0 likes33 downloads2y agoHugging Face26benjleite /FairytaleQA-translated-french Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the French machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-french.textquestion-answering10K<n<100K1 likes31 downloads1y agoHugging Face27CATIE-AQ /french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review Summary french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review is a subset of the Dataset of French Prompts (DFP).It contains 347,688 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset french_book_reviews.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_book_reviews_fr_prompt_binary_text_generation_from_title_of_a_review.texttext-generation100K<n<1M0 likes30 downloads1y agoHugging Face28CATIE-AQ /squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question Summary squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question is a subset of the Dataset of French Prompts (DFP).It contains 1,271,928 rows that can be used for a context-generation (with answer and question) task.The original data (without prompts) comes from the dataset pragnakalp/squad_v2_french_translated and was augmented by questions in SQUAD 2.0 format in the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question.texttext-generation1M<n<10M0 likes30 downloads1y agoHugging Face29VinceGx33 /mistral-legal-french-dataset Mistral Legal French Dataset A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy. 📋 Table of Contents Overview Dataset Composition Methodology 1. Chain-of-Thought Generation 2. LegalKit Extraction 3. Curriculum Learning Fusion Data Format Quality Metrics Usage Citations License 🎯 Overview This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two complementary… See the full description on the dataset page: https://huggingface.co/datasets/VinceGx33/mistral-legal-french-dataset.texttext-generation10K<n<100K2 likes30 downloads11mo agoHugging Face30finaleads /french-corpus-llm-sample French Corpus LLM — Sample 500 (v1.4.0) FINALEADS LLC builds compliance-ready training datasets for French regulated industries. We turn 2.66 billion tokens of French finance, regulatory, and economic open data into audit-trailed, pseudonymized, AI Act Article 10-documented shares — so foundation model and regtech teams can ship into European enterprises without a data-lineage gap. This is a public sample of 500 stratified documents drawn from the French Premium Web Corpus v1.4.0… See the full description on the dataset page: https://huggingface.co/datasets/finaleads/french-corpus-llm-sample.documenttext-generationn<1K1 likes30 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.