CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01singletongue /cc100-documents cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance corresponds to a single paragraph (or a document boundary). In this version, the data has been reformed so that each instance corresponds to a single, complete document. This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library. Languages The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.tabulartext-generation100M<n<1B1 likes1.1k downloads11mo agoHugging Face02th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M45 likes1.1k downloads2mo agoHugging Face03jinaai /wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.image1K<n<10K0 likes827 downloads1y agoHugging Face04hotchpotch /cc100-ja-documents cc100-ja-documents HuggingFace で公開されている cc100 / cc100-ja は line 単位の分割のため、document 単位に結合したものです。 ライセンスはオリジナルのcc100 に準拠します。 text10M<n<100M4 likes785 downloads2y agoHugging Face05DBWBD /Chinese_Debate_Documents Dataset Card for Chinese Debate Documents ASR-transcribed corpus of competitive Mandarin Chinese university debates, speaker-segmented and timestamped, with topic / round / team metadata parsed from the source filenames. Loading the Dataset from datasets import load_dataset ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train") print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"]) for seg in ds[0]["segments"][:3]: print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.tabulartext-classification1K<n<10K1 likes727 downloads4mo agoHugging Face06hulk10 /legi-full-documentstext100K<n<1M0 likes685 downloads6h agoHugging Face07hulk10 /data_gouv_datasets_catalog-full-documents 🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data. Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes : titre et description, organisation productrice, licence, couverture spatiale et temporelle, fréquence de mise à jour, formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.tabular100K<n<1M0 likes511 downloads6h agoHugging Face08hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes488 downloads6h agoHugging Face09hulk10 /conseil-administratives-appel-full-documents Décisions de Justice Administrative Françaises Description du dataset Ce dataset contient un corpus de décisions de justice administrative françaises extraites de la plateforme officielle Open Data de la Justice Administrative. Les documents sont publiés par le Conseil d'État dans le cadre de la politique d'ouverture des données publiques de la justice française. Les décisions sont diffusées au format XML et anonymisées conformément aux exigences légales relatives… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-administratives-appel-full-documents.texttext-classification1K<n<10K1 likes485 downloads6h agoHugging Face10hulk10 /local_administrations_directory-full-documents 🇫🇷 Référentiel des administrations locales – Version structurée Ce dataset regroupe l’Annuaire de l’administration – Base de données locales, qui recense l’ensemble des administrations et services publics locaux français : collectivités territoriales, services municipaux, services départementaux et régionaux, établissements publics locaux, structures administratives de proximité. Les données sont issues des sources open data officielles publiées sur data.gouv.fr et… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/local_administrations_directory-full-documents.text10K<n<100K1 likes398 downloads1d agoHugging Face11KanoonGPT /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.texttext-generation10K<n<100K0 likes384 downloads4mo agoHugging Face12hulk10 /service_public_part-full-documents 🇫🇷 Dataset Service-Public.fr – Fiches administratives structurées Ce dataset est constitué à partir des contenus officiels publiés sur la plateformeService-Public.fr.Il regroupe des fiches pratiques et ressources administratives à destination des particuliers et des professionnels, couvrant un large éventail de démarches et de thématiques de l’administration française. La structure et la méthodologie de ce dataset sont fortement inspirées du dataset Service-Public.fr practical… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/service_public_part-full-documents.textquestion-answering1K<n<10K0 likes381 downloads6h agoHugging Face13hulk10 /service_public_pro-full-documentstextquestion-answering1K<n<10K0 likes381 downloads6h agoHugging Face14hulk10 /travail_emploi-full-documents 🇫🇷 Dataset Ministère du Travail et de l’Emploi – Fiches structurées Ce dataset est constitué à partir des contenus publics diffusés sur le site officiel duMinistère du Travail et de l’Emploi :https://travail-emploi.gouv.fr/ Les données sources proviennent du dépôt GitHub officiel de l’administration française :https://github.com/SocialGouv/fiches-travail-data La structure et la logique générale de ce dataset sont inspirées du dataset Travail Emploi website Dataset, publié sur… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/travail_emploi-full-documents.textquestion-answeringn<1K0 likes358 downloads6h agoHugging Face15vohuutridung /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 2,393 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/vohuutridung/vietnamese-legal-documents.texttext-classification1M<n<10M3 likes328 downloads6mo agoHugging Face16geodesic-research /midtraining_mix_modernbert_filtered_documentstext1M<n<10M0 likes319 downloads10mo agoHugging Face17hulk10 /state_administrations_directory-full-documents 🇫🇷 Référentiel des administrations de l’État – Version structurée Ce dataset regroupe le Référentiel de l’organisation administrative de l’État, publié par la Direction de l’information légale et administrative (DILA).Il recense l’ensemble des administrations, services et organismes de l’État français, avec leurs missions, coordonnées, responsables et relations hiérarchiques. Les données sont issues des sources open data officielles : du portail data.gouv.fr, et du site… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/state_administrations_directory-full-documents.text1K<n<10K1 likes290 downloads6h agoHugging Face18hulk10 /legi-raw-documentstext1M<n<10M0 likes290 downloads9h agoHugging Face19hulk10 /tribunal-administratif-full-documents Tribunal Administratif - Full Documents Description Ce dataset contient un corpus de décisions issues des Tribunaux Administratifs français, converties en documents textuels exploitables pour les applications d'intelligence artificielle. L'objectif est de fournir un corpus prêt à l'emploi pour : le Retrieval-Augmented Generation (RAG) ; la recherche juridique ; la question-réponse ; la classification documentaire ; le fine-tuning de modèles de langage spécialisés… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/tribunal-administratif-full-documents.textquestion-answering10K<n<100K0 likes289 downloads6h agoHugging Face20jinaai /wikimedia-commons-documents-ml_deprecated Wikimedia Commons Document Retrieval Wikimedia Commons Documents This dataset is created for the evaluation of retrieval models. It contains images of (mostly historic) documents which should be identified based on their description. We extracted those descriptions from Wikimedia Commons. We have included the license type and a link (license_text) to the original Wikimedia Commons page for each extracted image. The text_description column contains OCR text extracted from the images… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_deprecated.image10K<n<100K1 likes285 downloads1y agoHugging Face21hulk10 /dole-raw-documentstext1K<n<10K0 likes259 downloads9h agoHugging Face22hf-internal-testing /example-documentsimagen<1K1 likes247 downloads2y agoHugging Face23sheggle /dutch-legal-documents Dutch Legal Documents A comprehensive collection of 1.14 million Dutch legal documents, including court rulings, parliamentary documents, and EU legislation. Sources Source Type Documents Description Rechtspraak.nl Court rulings 882,212 All Dutch court rulings from Hoge Raad, Raad van State, Gerechtshoven, Rechtbanken, CRvB, CBb Officiële Bekendmakingen Parliamentary docs 255,837 Kamerstukken, Handelingen, Memories van Toelichting, Staatsblad EUR-Lex EU… See the full description on the dataset page: https://huggingface.co/datasets/sheggle/dutch-legal-documents.tabular1M<n<10M1 likes238 downloads6mo agoHugging Face24demokratis /consultation-documents🚀 Demokratis.ch makes it easier to participate in Swiss consultation procedures in order to better influence the legislative process at the federal and cantonal level. What data we use We obtain information about federal and cantonal consultations through APIs and website scraping. For each consultation (Vernehmlassung) we typically collect a number of documents of various types: The proposed law change (draft, "Vorlage", "Entwurf", ...) A report explaining the proposed change… See the full description on the dataset page: https://huggingface.co/datasets/demokratis/consultation-documents.text100K<n<1M1 likes237 downloads3d agoHugging Face25hulk10 /dole-full-documents 🇫🇷 Dataset Dossiers législatifs (DOLE) – Version structurée Ce dataset regroupe les dossiers législatifs (DOLE) publiés par les institutions françaises.Il couvre l’ensemble des lois promulguées depuis la XIIᵉ législature (juin 2002), ainsi que : les projets de loi, les propositions de loi, les ordonnances, et les dossiers législatifs en cours d’élaboration. Les données sont issues des sources open data officielles mises à disposition par la DILA et référencées sur… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/dole-full-documents.text10K<n<100K0 likes231 downloads6h agoHugging Face26erdem-erdem /Turkish-Law-Documents-700k-clustered Turkish Legal Documents Clustering Dataset A comprehensive dataset of 700,000 Turkish legal documents from the two primary sources of legal precedent in Turkey, clustered using multiple emebdding models and algorithms to enable research, analysis, and machine learning applications. Overview This repository contains a large-scale document clustering pipeline and dataset for Turkish legal documents sourced from: Yargıtay - Turkey's highest court of appeal for civil and… See the full description on the dataset page: https://huggingface.co/datasets/erdem-erdem/Turkish-Law-Documents-700k-clustered.tabular100K<n<1M7 likes202 downloads11mo agoHugging Face27YuITC /Vietnamese-Legal-Documents Vietnamese Legal Documents Dataset 1. Dataset Summary Raw data: tmnam20/BKAI-Legal-Retrieval The Vietnamese Legal Documents Dataset is a benchmark dataset designed for legal information retrieval in the Vietnamese language. It consists of: A corpus of legal documents. Train/test splits containing natural language queries and their corresponding relevant documents. This dataset is intended to support research and development in: Information Retrieval (IR)… See the full description on the dataset page: https://huggingface.co/datasets/YuITC/Vietnamese-Legal-Documents.texttext-retrieval100K<n<1M6 likes197 downloads6mo agoHugging Face28hulk10 /cnil-raw-documentstext10K<n<100K0 likes195 downloads9h agoHugging Face29prg-unibe /dodis-historical-documents Dataset Overview The Swiss Historical Archive of DOdis Works (SHADOW) is a large-scale benchmark dataset derived from the resources of Dodis, an independent research center dedicated to the study of Swiss foreign policy and Switzerland’s international relations. Dodis has curated and published nearly 50,000 key documents that document the administrative practices and decision-making processes of the Swiss federal administration. The underlying corpus consists of notes, letters… See the full description on the dataset page: https://huggingface.co/datasets/prg-unibe/dodis-historical-documents.tabularsummarization10K<n<100K4 likes181 downloads4mo agoHugging Face30hulk10 /cnil-full-documents 🇫🇷 Dataset CNIL – Délibérations et décisions structurées Ce dataset regroupe les délibérations, décisions et actes officiels publiés par laCommission Nationale de l’Informatique et des Libertés (CNIL), autorité administrative indépendante chargée de la protection des données personnelles en France. Les données sont issues des sources open data officielles mises à disposition par la DILA et référencées sur data.gouv.fr.Elles couvrent un large spectre d’actes juridiques :… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/cnil-full-documents.text10K<n<100K0 likes159 downloads6h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.