datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scipar_parallel_docs
SciPar Parallel Documents
Dataset Description
This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts.
In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories.
This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences.
To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.real_clinical_cases_of_Famous_Old_TCM_Doctors
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors 数据集简介
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors是一个包含了当代著名老中医临床病例的数据集。这些病例数据来源于《当代名老中医典型医案集》(Contemporary Famous Old Chinese Medicine Doctors' Typical Cases Collection)一书。该数据集收录了多位德高望重的老中医大家的真实门诊病历,涵盖了多种常见病和疑难杂症。每个病例都包括病情描述、辨证论治思路、具体治疗方药等宝贵的一手临床资料。这些医案凝聚了老一辈名医的智慧和经验,对于中医的传承发展和临床应用研究,都有重要价值。通过对这些案例的挖掘分析,能够总结老中医诊疗思维、理法方药的特点,为现代中医临床实践提供有益借鉴。
Introduction to TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors… See the full description on the dataset page: https://huggingface.co/datasets/TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors.mozart-api-demo-pages
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.vn-provinces-doctors
Vietnam doctors by locality
Vietnam number of doctors by province/region, 2018-2022. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (315 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (30 rows)
data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-doctors.dochopcrh-parallel-corpora-document-level-noisydocling-nlp-datasetsThis repository contains the models used for docling-nlp.
Contents
This model repository packages the pretrained assets used by Docling’s NLP
components:
CRF models for material classification and English part-of-speech tagging
fastText models for language detection, metadata, semantic, topic, and person-name classification
Regular-expression assets for geographic-location extraction and unit handling
A default tokenizer model
Correct workflow to add new files… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/docling-nlp-datasets.legal-privilege-log-document-basis-waiver-risk-v0.1What this dataset does
You receive
doc description
date
author
recipients
privilege basis
redaction choice
context
waiver flags
You decide
coherent
or
incoherent
Daily use
privilege log QC
waiver risk detection
disclosure challenge prep
idris_doc_pairs_datasetBAREC-Shared-Task-2025-doc
BAREC Shared Task 2025
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-doc.linkedin_job_listingsdigits
Dataset Card for digits dataset
Optical recognition of handwritten digits dataset
Note - How to load this dataset directly with the datasets library
from datasets import load_dataset
dataset = load_dataset("sklearn-docs/digits",header=None)
Dataset Summary
This is a copy of the test set of the UCI ML hand-written digits datasets https://archive.ics.uci.edu/ml/datasets/Optical+Recognition+of+Handwritten+Digits
The data set contains images of hand-written… See the full description on the dataset page: https://huggingface.co/datasets/sklearn-docs/digits.document-photo-requirements
Verified Document Photo Requirements Dataset
A structured reference collection of official-source passport, visa, and national ID photo requirements maintained by Passport Visa Photo Studio. It is not a photo corpus, training dataset, or model artifact.
Dataset summary
Version: 1.0.0
Release date: 2026-08-14
Latest source review represented: 2026-08-11
Records: 18 (12 passport, 5 visa, 1 national ID)
Coverage: 14 countries or regions
Formats: CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/passport-visa-photo-studio/document-photo-requirements.DEplain-web-doc
DEplain-web-doc: A corpus for German Document Simplification
DEplain-web-doc is a subcorpus of DEplain Stodden et al., 2023 for document simplification.
The corpus consists of 396 (199/50/147) parallel documents crawled from the web in standard German and plain German (or easy-to-read German). All documents are either published under an open license or the copyright holders gave us the permission to share the data.
If you are interested in a larger corpus, please check our paper… See the full description on the dataset page: https://huggingface.co/datasets/DEplain/DEplain-web-doc.docs-python-v1
Dataset Card for Dataset Name
This dataset card aims to be a base template for creating python docs from methods. This is formatted from semeru/code-code-galeras-code-completion-from-docstring-3k-deduped
Dataset Description
Curated by: semeru/code-code-galeras-code-completion-from-docstring-3k-deduped
Language(s) (NLP): Python
License: [More Information Needed]
Dataset Sources [optional]
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/ASHu2/docs-python-v1.Multi-Doc-Multi-QA-ChineseDeprecated, please use Multi-Doc-QA-Chinese instead.
文档和问答对都来自 Multi-Doc-QA-Chinese,通过随机抽取和组合形成多轮问答形式。
推荐直接使用原始数据集Multi-Doc-QA-Chinese自己生成指令微调数据,可以控制参考文档和问答的数量
经过随机组合,每条数据形成了 20-60个参考文档 + 10个问答对的形式
chat格式为chatml
legal-doctrine-evolution-coherence-trajectory-v0.1What this dataset is
You receive
doctrine state at t
transition signals
doctrine state at t+1
split or exception signals
workability or legitimacy signals
reform pressure signals
You decide
Is the doctrine evolution stable
Answer
coherent
or
incoherent
Why this matters
Incoherent trajectories predict
overruling events
doctrinal collapse
rapid rule change
institutional instability
filipino-doctorThis dataset contains list of probing questions of doctors in Filipino
compile-benchmarksexperts-backendsDocVQA_valDocVQA_val_questions_completion_file_100legal-document-version-redline-final-coherence-risk-v0.1What this dataset does
You receive
version history
redline summary
final id
sent or filed id
approval record
mismatch flags
You decide
coherent
or
incoherent
Daily use
wrong attachment prevention
filing version QC
approval gap detection
DocVQA_IncompleteDocVQA_Question_Completion_valscotus-docket-dataset-rawlegal-chronology-event-document-issue-coherence-risk-v0.1What this dataset does
You receive
timeline summary
document map
issue links
date checks
gap flags
conflict flags
You decide
coherent
or
incoherent
Daily use
chronology QC
date conflict detection
missing evidence detection
gap finding
DocVQA_VALVLMEvalKit_sourced_DocVQA_valDocVQA_100_QC_val
