datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
doc-formats-csv-1
[doc] formats - csv - 1
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
The YAML section of the README does not contain anything related to loading the data (only the size category metadata):
---
size_categories:
- n<1K
---
huggingface_doczamai-pashto-documents
ZamAI-Pashto Documents
This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language.
Project Structure
data/: Contains scanned documents, extracted text, translations, and summaries.
annotations/: OCR bounding boxes, handwriting labels, and domain tags.
scripts/: OCR processing, text cleaning, and translation alignment scripts.
configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.Patient-Doctor-Conversationcavl-doc-lacdip-split3
CaVL-Doc — LA-CDIP Final Split 3 (Augmented)
Dataset de treino e validação para o modelo definitivo CaVL-Doc, baseado no
Split 3 do protocolo ZSL (Zero-Shot Learning) sobre o LA-CDIP.
Estrutura
Conjunto
Imagens
Pares
Treino (images_train/)
8,740 variantes augmentadas
34.960
Validação (images_val/)
2,625 variantes augmentadas
10.500
120 classes para treino · 24 classes novel para validação (sem sobreposição)
Cada imagem original gera 5 variantes com o… See the full description on the dataset page: https://huggingface.co/datasets/Jpcosta90/cavl-doc-lacdip-split3.scipar_parallel_docs
SciPar Parallel Documents
Dataset Description
This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts.
In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories.
This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences.
To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.real_clinical_cases_of_Famous_Old_TCM_Doctors
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors 数据集简介
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors是一个包含了当代著名老中医临床病例的数据集。这些病例数据来源于《当代名老中医典型医案集》(Contemporary Famous Old Chinese Medicine Doctors' Typical Cases Collection)一书。该数据集收录了多位德高望重的老中医大家的真实门诊病历,涵盖了多种常见病和疑难杂症。每个病例都包括病情描述、辨证论治思路、具体治疗方药等宝贵的一手临床资料。这些医案凝聚了老一辈名医的智慧和经验,对于中医的传承发展和临床应用研究,都有重要价值。通过对这些案例的挖掘分析,能够总结老中医诊疗思维、理法方药的特点,为现代中医临床实践提供有益借鉴。
Introduction to TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors… See the full description on the dataset page: https://huggingface.co/datasets/TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors.qdrant_docmozart-api-demo-pages
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.patient-doctor-qa-tr-321179
Patient Doctor Q&A TR 321179 Veri Kümesi
Patient Doctor Q&A TR 321179 veri kümesi, Patient Doctor Q&A TR 19583, Patient Doctor Q&A TR 167732, Patient Doctor Q&A TR 5695 ve Patient Doctor Q&A TR 95588 veri kümelerinin birleştirilmiş ve karıştırılmış halidir.
Ana Özellikler:
İçerik: Çeşitli tıbbi konuları kapsayan hasta soruları ve doktor yanıtları.
Yapı: 2 sütun içerir: Soru, Cevap.Dil: Türkçe.
Potansiyel Kullanım Alanları:
Tıbbi araştırmalar
Doğal Dil… See the full description on the dataset page: https://huggingface.co/datasets/kayrab/patient-doctor-qa-tr-321179.vn-provinces-doctors
Vietnam doctors by locality
Vietnam number of doctors by province/region, 2018-2022. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (315 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (30 rows)
data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-doctors.streamlit_docsdochoplegal_consolidationTask details
Legal consolidation is a critical yet time-consuming task, traditionally performed manually by legal professionals.
The objective is to automate the process of French legal consolidation, which is the application of modifications from a
modification section to an initial article to generate a modified article.
Dataset structure
A triplet of:
an initial article: the legislative article before consolidation,
a modification section: the text introducing the modification within… See the full description on the dataset page: https://huggingface.co/datasets/DoctrineAI/legal_consolidation.qdrant_doc_qnamozart-api
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api.easylaw_kr_documentsFrench_Doctoral_Theses
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/antoinebourgois2/french-doctoral-thesis
Description
All french doctoral thesis metatdata scrapped from https://www.theses.fr
The dataset contains :
URL
Thesis title
Short description
Author name
Thesis director(s) informations
Research domain
Status ( defended / in preparation)
doc-splits-1
[doc] file names and splits 1
This dataset contains a data.csv file at the root.
MNLP_M3_rag_documentscrh-parallel-corpora-document-level-noisydocling-nlp-datasetsThis repository contains the models used for docling-nlp.
Contents
This model repository packages the pretrained assets used by Docling’s NLP
components:
CRF models for material classification and English part-of-speech tagging
fastText models for language detection, metadata, semantic, topic, and person-name classification
Regular-expression assets for geographic-location extraction and unit handling
A default tokenizer model
Correct workflow to add new files… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/docling-nlp-datasets.legal-privilege-log-document-basis-waiver-risk-v0.1What this dataset does
You receive
doc description
date
author
recipients
privilege basis
redaction choice
context
waiver flags
You decide
coherent
or
incoherent
Daily use
privilege log QC
waiver risk detection
disclosure challenge prep
maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for
Triage trade doc packs before they trigger holds.
You use it to flag
HS code inconsistencies across documents
missing certificates
shipper or consignee mismatch
clearance status lag not supported by doc quality
Why it matters
Most port delay disputes begin in paperwork.
idris_doc_pairs_datasetdatasette-spike-fara
Datasette spike — FARA Active Foreign Principals
For: CoS → WordPress Guru (doctorparadox.net embed/link)Built: 2026-09-17 (ET)Status: DATA half ready — public SQLite + Datasette Lite URL
Why this dataset
Doctor Paradox already centers corruption / foreign influence / authoritarian-adjacent reporting (Corruption Tracker, Corruption Daily cards). FARA filings are the federal public ledger of who lobbies in the U.S. on behalf of foreign principals.
We use the… See the full description on the dataset page: https://huggingface.co/datasets/doctorparadox/datasette-spike-fara.BAREC-Shared-Task-2025-doc
BAREC Shared Task 2025
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-doc.doc-splits-3
[doc] file names and splits 3
This dataset contains three csv files at the root: my_train_file.csv, test-file.csv, validation1.csv.
patient-doctor-qa-tr-95588
Patient Doctor Q&A TR 95588 Veri Seti
Patient Doctor Q&A TR 95588 veri seti, chat_doctor veri setinin Türkçeye çevrilmiş halidir.
Ana Özellikler:
İçerik: Çeşitli tıbbi konuları kapsayan hasta soruları ve doktor yanıtları.
Yapı: 3 sütun içerir: Talimat, Soru, Cevap.
Dil: Türkçe.
Potansiyel Kullanım Alanları:
Tıbbi araştırmalar
Doğal Dil İşleme (NLP)
Tıbbi eğitim
Sınırlamalar:
Veri gizliliği endişeleri
Yanıt kalitesinde değişkenlik
Potansiyel… See the full description on the dataset page: https://huggingface.co/datasets/kayrab/patient-doctor-qa-tr-95588.docids
