CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01datasets-examples /doc-formats-csv-1 [doc] formats - csv - 1 This dataset contains one csv file at the root: data.csv kind,sound dog,woof cat,meow pokemon,pika human,hello The YAML section of the README does not contain anything related to loading the data (only the size category metadata): --- size_categories: - n<1K --- textn<1K0 likes1.9k downloads3y agoHugging Face02m-ric /huggingface_doctext1K<n<10K21 likes1.4k downloads3y agoHugging Face03ZamAI-Pashto /zamai-pashto-documents ZamAI-Pashto Documents This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language. Project Structure data/: Contains scanned documents, extracted text, translations, and summaries. annotations/: OCR bounding boxes, handwriting labels, and domain tags. scripts/: OCR processing, text cleaning, and translation alignment scripts. configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.textvisual-document-retrievaln<1K0 likes612 downloads2mo agoHugging Face04mahfoos /Patient-Doctor-Conversationtext1K<n<10K17 likes423 downloads3y agoHugging Face05Jpcosta90 /cavl-doc-lacdip-split3 CaVL-Doc — LA-CDIP Final Split 3 (Augmented) Dataset de treino e validação para o modelo definitivo CaVL-Doc, baseado no Split 3 do protocolo ZSL (Zero-Shot Learning) sobre o LA-CDIP. Estrutura Conjunto Imagens Pares Treino (images_train/) 8,740 variantes augmentadas 34.960 Validação (images_val/) 2,625 variantes augmentadas 10.500 120 classes para treino · 24 classes novel para validação (sem sobreposição) Cada imagem original gera 5 variantes com o… See the full description on the dataset page: https://huggingface.co/datasets/Jpcosta90/cavl-doc-lacdip-split3.textimage-to-image10K<n<100K0 likes230 downloads4mo agoHugging Face06ilsp /scipar_parallel_docs SciPar Parallel Documents Dataset Description This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts. In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories. This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences. To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.tabulartext-generation1K<n<10K1 likes188 downloads2y agoHugging Face07TCMLM /real_clinical_cases_of_Famous_Old_TCM_Doctors TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors 数据集简介 TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors是一个包含了当代著名老中医临床病例的数据集。这些病例数据来源于《当代名老中医典型医案集》(Contemporary Famous Old Chinese Medicine Doctors' Typical Cases Collection)一书。该数据集收录了多位德高望重的老中医大家的真实门诊病历,涵盖了多种常见病和疑难杂症。每个病例都包括病情描述、辨证论治思路、具体治疗方药等宝贵的一手临床资料。这些医案凝聚了老一辈名医的智慧和经验,对于中医的传承发展和临床应用研究,都有重要价值。通过对这些案例的挖掘分析,能够总结老中医诊疗思维、理法方药的特点,为现代中医临床实践提供有益借鉴。 Introduction to TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors… See the full description on the dataset page: https://huggingface.co/datasets/TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors.tabularn<1K2 likes134 downloads3mo agoHugging Face08atitaarora /qdrant_doctextquestion-answeringn<1K0 likes123 downloads2y agoHugging Face09DoctorSlimm /mozart-api-demo-pages Dataset Card for Dataset Name Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.imagen<1K0 likes115 downloads3y agoHugging Face10kayrab /patient-doctor-qa-tr-321179 Patient Doctor Q&A TR 321179 Veri Kümesi Patient Doctor Q&A TR 321179 veri kümesi, Patient Doctor Q&A TR 19583, Patient Doctor Q&A TR 167732, Patient Doctor Q&A TR 5695 ve Patient Doctor Q&A TR 95588 veri kümelerinin birleştirilmiş ve karıştırılmış halidir. Ana Özellikler: İçerik: Çeşitli tıbbi konuları kapsayan hasta soruları ve doktor yanıtları. Yapı: 2 sütun içerir: Soru, Cevap.Dil: Türkçe. Potansiyel Kullanım Alanları: Tıbbi araştırmalar Doğal Dil… See the full description on the dataset page: https://huggingface.co/datasets/kayrab/patient-doctor-qa-tr-321179.textquestion-answering100K<n<1M5 likes112 downloads2y agoHugging Face11letrinhan /vn-provinces-doctors Vietnam doctors by locality Vietnam number of doctors by province/region, 2018-2022. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files provinces (315 rows) data/provinces.csv data/provinces.dta data/provinces.xlsx regions (30 rows) data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-doctors.tabularn<1K0 likes100 downloads1d agoHugging Face12sai-lohith /streamlit_docstextn<1K0 likes88 downloads2y agoHugging Face13zhuoranyu336 /dochoptabular1K<n<10K1 likes83 downloads2mo agoHugging Face14DoctrineAI /legal_consolidationTask details Legal consolidation is a critical yet time-consuming task, traditionally performed manually by legal professionals. The objective is to automate the process of French legal consolidation, which is the application of modifications from a modification section to an initial article to generate a modified article. Dataset structure A triplet of: an initial article: the legislative article before consolidation, a modification section: the text introducing the modification within… See the full description on the dataset page: https://huggingface.co/datasets/DoctrineAI/legal_consolidation.text1K<n<10K1 likes78 downloads3y agoHugging Face15atitaarora /qdrant_doc_qnatextn<1K1 likes64 downloads2y agoHugging Face16DoctorSlimm /mozart-api Dataset Card for Dataset Name Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api.imagen<1K0 likes60 downloads3y agoHugging Face17jiwoochris /easylaw_kr_documentstext1K<n<10K2 likes58 downloads3y agoHugging Face18FrancophonIA /French_Doctoral_Theses [!NOTE] Dataset origin: https://www.kaggle.com/datasets/antoinebourgois2/french-doctoral-thesis Description All french doctoral thesis metatdata scrapped from https://www.theses.fr The dataset contains : URL Thesis title Short description Author name Thesis director(s) informations Research domain Status ( defended / in preparation) text100K<n<1M1 likes56 downloads1y agoHugging Face19datasets-examples /doc-splits-1 [doc] file names and splits 1 This dataset contains a data.csv file at the root. textn<1K1 likes52 downloads3y agoHugging Face20flaviawallen /MNLP_M3_rag_documentstext1K<n<10K0 likes52 downloads1y agoHugging Face21QIRIM /crh-parallel-corpora-document-level-noisytabulartranslation10K<n<100K1 likes51 downloads2y agoHugging Face22docling-project /docling-nlp-datasetsThis repository contains the models used for docling-nlp. Contents This model repository packages the pretrained assets used by Docling’s NLP components: CRF models for material classification and English part-of-speech tagging fastText models for language detection, metadata, semantic, topic, and person-name classification Regular-expression assets for geographic-location extraction and unit handling A default tokenizer model Correct workflow to add new files… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/docling-nlp-datasets.tabular100K<n<1M0 likes51 downloads11d agoHugging Face23ClarusC64 /legal-privilege-log-document-basis-waiver-risk-v0.1What this dataset does You receive doc description date author recipients privilege basis redaction choice context waiver flags You decide coherent or incoherent Daily use privilege log QC waiver risk detection disclosure challenge prep tabulartext-classificationn<1K0 likes49 downloads7mo agoHugging Face24ClarusC64 /maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for Triage trade doc packs before they trigger holds. You use it to flag HS code inconsistencies across documents missing certificates shipper or consignee mismatch clearance status lag not supported by doc quality Why it matters Most port delay disputes begin in paperwork. texttext-classificationn<1K1 likes49 downloads7mo agoHugging Face25CNRS-IDRIS /idris_doc_pairs_datasettabular1K<n<10K0 likes47 downloads2y agoHugging Face26doctorparadox /datasette-spike-fara Datasette spike — FARA Active Foreign Principals For: CoS → WordPress Guru (doctorparadox.net embed/link)Built: 2026-09-17 (ET)Status: DATA half ready — public SQLite + Datasette Lite URL Why this dataset Doctor Paradox already centers corruption / foreign influence / authoritarian-adjacent reporting (Corruption Tracker, Corruption Daily cards). FARA filings are the federal public ledger of who lobbies in the U.S. on behalf of foreign principals. We use the… See the full description on the dataset page: https://huggingface.co/datasets/doctorparadox/datasette-spike-fara.textn<1K0 likes43 downloads5d agoHugging Face27CAMeL-Lab /BAREC-Shared-Task-2025-doc BAREC Shared Task 2025 Dataset Summary BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes. The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-doc.tabulartext-classification1K<n<10K2 likes42 downloads1y agoHugging Face28datasets-examples /doc-splits-3 [doc] file names and splits 3 This dataset contains three csv files at the root: my_train_file.csv, test-file.csv, validation1.csv. textn<1K0 likes41 downloads3y agoHugging Face29kayrab /patient-doctor-qa-tr-95588 Patient Doctor Q&A TR 95588 Veri Seti Patient Doctor Q&A TR 95588 veri seti, chat_doctor veri setinin Türkçeye çevrilmiş halidir. Ana Özellikler: İçerik: Çeşitli tıbbi konuları kapsayan hasta soruları ve doktor yanıtları. Yapı: 3 sütun içerir: Talimat, Soru, Cevap. Dil: Türkçe. Potansiyel Kullanım Alanları: Tıbbi araştırmalar Doğal Dil İşleme (NLP) Tıbbi eğitim Sınırlamalar: Veri gizliliği endişeleri Yanıt kalitesinde değişkenlik Potansiyel… See the full description on the dataset page: https://huggingface.co/datasets/kayrab/patient-doctor-qa-tr-95588.textquestion-answering100K<n<1M6 likes41 downloads2y agoHugging Face30alexia-allal /docidstext1M<n<10M0 likes41 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.