CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EleutherAI /wikitext_document_level Wikitext Document Level This is a modified version of https://huggingface.co/datasets/wikitext that returns Wiki pages instead of Wiki text line-by-line. The original readme is contained below. Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/wikitext_document_level.text10K<n<100K18 likes79k downloads2y agoHugging Face02HuggingFaceM4 /DocumentVQAimage10K<n<100K46 likes2.7k downloads3y agoHugging Face03RoboCOIN /Agilex_Cobot_Magic_zip_up_the_document_baggated Agilex_Cobot_Magic_zip_up_the_document_bag 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: Agilex_Cobot_Magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pull place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_zip_up_the_document_bag.tabularrobotics100K<n<1M0 likes1.3k downloads3mo agoHugging Face04th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M45 likes1.1k downloads2mo agoHugging Face05singletongue /cc100-documents cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance corresponds to a single paragraph (or a document boundary). In this version, the data has been reformed so that each instance corresponds to a single, complete document. This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library. Languages The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.tabulartext-generation100M<n<1B1 likes1k downloads11mo agoHugging Face06lingamvamshikrishnareddy /ramanv-document-ocr-2gatedtext100K<n<1M2 likes1k downloads18d agoHugging Face07RoboCOIN /RMC-AIDA-L_organise_the_document_baggated RMC-AIDA-L_organise_the_document_bag 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: realman_rmc_aidal | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp place pick pull 📊 Dataset Statistics Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_organise_the_document_bag.tabularrobotics100K<n<1M0 likes983 downloads9mo agoHugging Face08jinaai /wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.image1K<n<10K0 likes819 downloads1y agoHugging Face09DBWBD /Chinese_Debate_Documents Dataset Card for Chinese Debate Documents ASR-transcribed corpus of competitive Mandarin Chinese university debates, speaker-segmented and timestamped, with topic / round / team metadata parsed from the source filenames. Loading the Dataset from datasets import load_dataset ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train") print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"]) for seg in ds[0]["segments"][:3]: print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.tabulartext-classification1K<n<10K1 likes769 downloads4mo agoHugging Face10himalaya-ai /ocr-document-processing-eval ocr_document_processing_eval Document-processing OCR proxy for digitization, KYC-like numeric fields, and RAG extraction checks. Repo: himalaya-ai/ocr-document-processing-eval Task: document_processing_ocr Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns. Optional fine-tuning/eval file: *.sharegpt.json with messages and images. Core Columns id: unique sample identifier image: relative path to the image file ocr:… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/ocr-document-processing-eval.imageimage-to-textn<1K1 likes723 downloads3mo agoHugging Face11hotchpotch /cc100-ja-documents cc100-ja-documents HuggingFace で公開されている cc100 / cc100-ja は line 単位の分割のため、document 単位に結合したものです。 ライセンスはオリジナルのcc100 に準拠します。 text10M<n<100M4 likes680 downloads2y agoHugging Face12hulk10 /legi-full-documentstext100K<n<1M0 likes666 downloads3h agoHugging Face13hulk10 /data_gouv_datasets_catalog-full-documents 🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data. Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes : titre et description, organisation productrice, licence, couverture spatiale et temporelle, fréquence de mise à jour, formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.tabular100K<n<1M0 likes518 downloads3h agoHugging Face14hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes490 downloads4h agoHugging Face15hulk10 /conseil-administratives-appel-full-documents Décisions de Justice Administrative Françaises Description du dataset Ce dataset contient un corpus de décisions de justice administrative françaises extraites de la plateforme officielle Open Data de la Justice Administrative. Les documents sont publiés par le Conseil d'État dans le cadre de la politique d'ouverture des données publiques de la justice française. Les décisions sont diffusées au format XML et anonymisées conformément aux exigences légales relatives… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-administratives-appel-full-documents.texttext-classification1K<n<10K1 likes486 downloads4h agoHugging Face16Samoed /msmarco-documenttext1M<n<10M0 likes444 downloads10mo agoHugging Face17hulk10 /local_administrations_directory-full-documents 🇫🇷 Référentiel des administrations locales – Version structurée Ce dataset regroupe l’Annuaire de l’administration – Base de données locales, qui recense l’ensemble des administrations et services publics locaux français : collectivités territoriales, services municipaux, services départementaux et régionaux, établissements publics locaux, structures administratives de proximité. Les données sont issues des sources open data officielles publiées sur data.gouv.fr et… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/local_administrations_directory-full-documents.text10K<n<100K1 likes394 downloads3h agoHugging Face18hulk10 /service_public_part-full-documents 🇫🇷 Dataset Service-Public.fr – Fiches administratives structurées Ce dataset est constitué à partir des contenus officiels publiés sur la plateformeService-Public.fr.Il regroupe des fiches pratiques et ressources administratives à destination des particuliers et des professionnels, couvrant un large éventail de démarches et de thématiques de l’administration française. La structure et la méthodologie de ce dataset sont fortement inspirées du dataset Service-Public.fr practical… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/service_public_part-full-documents.textquestion-answering1K<n<10K0 likes389 downloads3h agoHugging Face19KanoonGPT /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.texttext-generation10K<n<100K0 likes386 downloads4mo agoHugging Face20hf-internal-testing /document-visual-retrieval-test Model Card: Document Visual Retrieval Test (internal) Dataset Overview This dataset is designed to evaluate the performance of visual retrievers by testing their ability to match a query to a relevant image. Each of the three examples in this dataset contains a text query and an associated image, which is a scanned page from the foundational "Attention is All You Need" paper. The purpose of this dataset is to facilitate the evaluation of visual retrievers, where the… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/document-visual-retrieval-test.imagen<1K1 likes375 downloads2y agoHugging Face21hulk10 /service_public_pro-full-documentstextquestion-answering1K<n<10K0 likes375 downloads3h agoHugging Face22ltg /norec_document Dataset Card for NoReC_document Document-level polarity classification of Norwegian full-text reviews across mixed domains. Dataset Details This is a dataset for document-level sentiment classification in Norwegian, derived from the Norwegian Review Corpus: NoReC. We here provide two simplified versions of NoReC where the original six-point numerical ratings have been mapped to a reduced set of categorical classes: positive and negative for the binary version, and… See the full description on the dataset page: https://huggingface.co/datasets/ltg/norec_document.texttext-classification10K<n<100K1 likes373 downloads2y agoHugging Face23hulk10 /travail_emploi-full-documents 🇫🇷 Dataset Ministère du Travail et de l’Emploi – Fiches structurées Ce dataset est constitué à partir des contenus publics diffusés sur le site officiel duMinistère du Travail et de l’Emploi :https://travail-emploi.gouv.fr/ Les données sources proviennent du dépôt GitHub officiel de l’administration française :https://github.com/SocialGouv/fiches-travail-data La structure et la logique générale de ce dataset sont inspirées du dataset Travail Emploi website Dataset, publié sur… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/travail_emploi-full-documents.textquestion-answeringn<1K0 likes357 downloads3h agoHugging Face24deeptools-ai /test-document-invoiceimagen<1K2 likes352 downloads4y agoHugging Face25biglam /icdar2021-historical-document-dating ICDAR 2021 Historical Document Classification — Task 2 (Dating) 13,810 manuscript page images labelled with the period in which they were produced. Images come from e-codices, the virtual manuscript library of Switzerland. Split Images Date range Median span Dated to a single year train 11,294 800–1899 45 years 1,409 test 2,516 800–1921 49 years 264 The label is an interval, not a year Palaeographers date a manuscript to a range, and the width of… See the full description on the dataset page: https://huggingface.co/datasets/biglam/icdar2021-historical-document-dating.imageimage-classification10K<n<100K2 likes351 downloads2mo agoHugging Face26vohuutridung /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 2,393 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/vohuutridung/vietnamese-legal-documents.texttext-classification1M<n<10M3 likes334 downloads6mo agoHugging Face27geodesic-research /midtraining_mix_modernbert_filtered_documentstext1M<n<10M0 likes312 downloads10mo agoHugging Face28hulk10 /legi-raw-documentstext1M<n<10M0 likes303 downloads7h agoHugging Face29NeuralMetrics /DocumentVQA Neural Metrics · Asking documents questions, and grading the answers. Document visual question answering: given a page image and a natural-language question, produce the answer. It measures whether a model genuinely read the layout or merely pattern-matched the text. We use it for: evaluating question-answering over extracted documents - catching models that read text but misread structure. Attribution This is an unmodified fork of… See the full description on the dataset page: https://huggingface.co/datasets/NeuralMetrics/DocumentVQA.image10K<n<100K0 likes294 downloads2mo agoHugging Face30hulk10 /state_administrations_directory-full-documents 🇫🇷 Référentiel des administrations de l’État – Version structurée Ce dataset regroupe le Référentiel de l’organisation administrative de l’État, publié par la Direction de l’information légale et administrative (DILA).Il recense l’ensemble des administrations, services et organismes de l’État français, avec leurs missions, coordonnées, responsables et relations hiérarchiques. Les données sont issues des sources open data officielles : du portail data.gouv.fr, et du site… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/state_administrations_directory-full-documents.text1K<n<10K1 likes291 downloads3h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.