datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opm-ehri-dataehri-pl-lines
ehri-pl-lines
Line-level OCR/HTR dataset of Polish typewritten historical documents, derived
from the Polish sub-corpus of the EHRI dataset.
Each example is a single text-line crop plus its ground-truth transcription. Line
bounding boxes come from the original ALTO XML ground truth (no automatic
detection was used), so transcriptions are reliable and aligned.
Content and provenance
468 line crops from 15 pages across 6 documents (ZIH collection).
Source images… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/ehri-pl-lines.patent-desc-testAfrica-Blood-Cell-Images-and-EHR-for-Cancer-Detection
Africa Blood Cell Images and EHR for Cancer Detection | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-Blood-Cell-Images-and-EHR-for-Cancer-Detection.africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all
Brain Tumor (MRI) Detection Colourized with EHR | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: not declared - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all.ehri-danish-typewritten-lines
EHRI Danish Typewritten Lines
This dataset contains 1,007 cropped line images from 36 Danish typewritten pages in the EHRI Dataset. The source material consists of Danish diplomatic reports from the Second World War held by the Danish National Archives.
It is published as a ready-to-use line recognition evaluation dataset. It is typewritten material, not book or newspaper typesetting.
Dataset structure
The dataset has one test split because the upstream dataset… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/ehri-danish-typewritten-lines.synth-ehr-icd10cm-promptehristoforu__RQwen-v0.2-details
Dataset Card for Evaluation run of ehristoforu/RQwen-v0.2
Dataset automatically created during the evaluation run of model ehristoforu/RQwen-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__RQwen-v0.2-details.Africa-skin-cancer-images-EHR
Africa skin cancer images EHR | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet, optimized-parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-skin-cancer-images-EHR.EHR-Ins-Reasoning
EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis
This repository contains the EHR-Ins-Reasoning dataset, as presented in the paper EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis.
EHR-Ins is a large-scale, comprehensive instruction dataset developed to enhance the reasoning and analysis capabilities of Large Language Models (LLMs) for Electronic Health Records (EHR).
Composition and… See the full description on the dataset page: https://huggingface.co/datasets/BlueZeros/EHR-Ins-Reasoning.ehristoforu__della-70b-test-v1-details
Dataset Card for Evaluation run of ehristoforu/della-70b-test-v1
Dataset automatically created during the evaluation run of model ehristoforu/della-70b-test-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__della-70b-test-v1-details.HIV_EHR_dastaset
Dataset Description
This dataset is a large-scale collection of 516,691 HIV patient records, designed to support the development of advanced healthcare AI systems, medical analytics, clinical decision support tools, and healthcare research applications.
It consists of real-world HIV clinical records collected from healthcare and treatment environments, containing structured patient information related to disease monitoring, treatment management, laboratory measurements, clinical… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/HIV_EHR_dastaset.ehristoforu__SoRu-0009-details
Dataset Card for Evaluation run of ehristoforu/SoRu-0009
Dataset automatically created during the evaluation run of model ehristoforu/SoRu-0009
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__SoRu-0009-details.medical-ehr-training-data
Medical EHR Training Dataset
Training dataset for Medical EHR GEPA-optimized module.
Dataset Description
This dataset contains 382 medical EHR query examples for training DSPy GEPA optimization.
Dataset Structure
{
"query": "Show me diabetic patients",
"expected_strategy": "ENRICHMENT",
"expected_snomed_codes": ["73211009", "44054006"],
"expected_neo4j_count": 15,
"query_complexity": "simple",
"medical_category": "endocrine"
}
Splits… See the full description on the dataset page: https://huggingface.co/datasets/Fanoni/medical-ehr-training-data.ehristoforu__qwen2.5-test-32b-it-details
Dataset Card for Evaluation run of ehristoforu/qwen2.5-test-32b-it
Dataset automatically created during the evaluation run of model ehristoforu/qwen2.5-test-32b-it
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__qwen2.5-test-32b-it-details.asia-ilo-ear-ehra-sex-age-cur-nb-average-hourly-earnings-of-employees-by-sex-and-ag
Average hourly earnings of employees by sex and age and currency | Asia (ILOSTAT)
🌏 14,982 observations · 25 Asia countries · 1996–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 14,982 observations of Earnings data across 25 Asia countries, spanning 1996–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ear-ehra-sex-age-cur-nb-average-hourly-earnings-of-employees-by-sex-and-ag.ehrsql_mimic_iii
Dataset Card for "ehrsql_processed"
More Information needed
asia-ilo-ear-ehra-sex-ocu-cur-nb-average-hourly-earnings-of-employees-by-sex-occupa
Average hourly earnings of employees by sex, occupation and currency | Asia (ILOSTAT)
🌏 30,895 observations · 31 Asia countries · 1997–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 30,895 observations of Earnings data across 31 Asia countries, spanning 1997–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ear-ehra-sex-ocu-cur-nb-average-hourly-earnings-of-employees-by-sex-occupa.asia-ilo-ear-ehra-sex-edu-cur-nb-average-hourly-earnings-of-employees-by-sex-educat
Average hourly earnings of employees by sex, education and currency | Asia (ILOSTAT)
🌏 22,960 observations · 25 Asia countries · 1996–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 22,960 observations of Earnings data across 25 Asia countries, spanning 1996–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ear-ehra-sex-edu-cur-nb-average-hourly-earnings-of-employees-by-sex-educat.asia-ilo-ear-ehra-sex-ind-cur-nb-average-hourly-earnings-of-employees-by-ilo-sector
Average hourly earnings of employees by ILO sector, sex and currency | Asia (ILOSTAT)
🌏 29,390 observations · 22 Asia countries · 2005–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 29,390 observations of Earnings data across 22 Asia countries, spanning 2005–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ear-ehra-sex-ind-cur-nb-average-hourly-earnings-of-employees-by-ilo-sector.synth-ehr-icd10-alpaca-formatehr_relEHR-Rel is a novel open-source1 biomedical concept relatedness dataset consisting of 3630 concept pairs, six times more
than the largest existing dataset. Instead of manually selecting and pairing concepts as done in previous work,
the dataset is sampled from EHRs to ensure concepts are relevant for the EHR concept retrieval task.
A detailed analysis of the concepts in the dataset reveals a far larger coverage compared to existing datasets.mathing-1
Dataset Card for "mathing-1"
More Information needed
Synthetic-EHR-Mistralasia-ilo-ear-ehra-sex-eco-cur-nb-average-hourly-earnings-of-employees-by-sex-econom
Average hourly earnings of employees by sex, economic activity and currency | Asia (ILOSTAT)
🌏 54,975 observations · 25 Asia countries · 1996–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 54,975 observations of Earnings data across 25 Asia countries, spanning 1996–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ear-ehra-sex-eco-cur-nb-average-hourly-earnings-of-employees-by-sex-econom.Synthetic-EHR-LlamaDeidentification-of-EHRsynth-ehr-icd10-llama3-formatchunked-ehrsynthetic-medical-ehr-dataset
🏥 Synthetic Privacy-Preserving Medical EHR Dataset
Dataset Description
A fully synthetic collection of 10,000 Electronic Health Records (EHRs) for binary classification research. The task is predicting adverse patient outcomes — deterioration or death — from clinical, demographic, and vital sign features. No real patient data was used at any stage. Privacy-safe under GDPR, CCPA, and HIPAA.
Supported Tasks
tabular-classification: Predict adverse_outcome (0/1)… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/synthetic-medical-ehr-dataset.
