datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ECGBench
ECGBench
Benchmark Dataset for paper "Teach Multimodal LLMs to Comprehend Electrocardiographic Images".
🌐 Project Page: https://aimedlab.github.io/PULSE/
📄 Paper: https://arxiv.org/abs/2410.19008
🧑💻 Code: https://github.com/AIMedLab/PULSE
🤗 Model: https://huggingface.co/PULSE-ECG/PULSE-7B
👩⚕️ ECGInstruct: https://huggingface.co/datasets/PULSE-ECG/ECGInstruct
Introduction
We introduce ECGBench, a comprehensive benchmark designed to evaluate ECG image… See the full description on the dataset page: https://huggingface.co/datasets/PULSE-ECG/ECGBench.PulseBench-Tab
PulseBench-Tab
A frontier multilingual benchmark for table extraction from document images.
PulseBench-Tab contains 1,820 human-annotated tables across 9 languages (English 589, Chinese 213, Spanish 176, Russian 170, French 165, Japanese 164, Arabic 146, German 113, Korean 84) and 4 scripts (Latin, CJK, Arabic, Cyrillic), sourced from 380 unique documents including financial filings, government reports, corporate disclosures, and regulatory filings. Each sample is a table image… See the full description on the dataset page: https://huggingface.co/datasets/pulse-ai/PulseBench-Tab.pulse-ecg-instruct-subsetlondons-pulse-moh
London's Pulse: Medical Officer of Health reports (page images + OCR text)
Page-level scans of the Wellcome Collection London's Pulse
Medical Officer of Health (MOH) reports (1848–1972), paired with OCR text, per-report
licence, and full provenance. Built for OCR / VLM / document-understanding work on real
historical public-health records — dense statistical tables, mixed layouts, century-old print.
Configs
config
rows
what
default
391,964 pages / 4,886… See the full description on the dataset page: https://huggingface.co/datasets/biglam/londons-pulse-moh.PulseBench-Select
PulseBench-Select
A benchmark for selected-option detection in document images.
PulseBench-Select contains 485 cleaned document images with ground-truth annotations for checkboxes, radio buttons, and marked answer choices. Each sample pairs a document image with public ground truth for the visible options that are selected.
Scoring methodology (GitHub): https://github.com/Pulse-Software-Corp/PulseBench-Select
Quick Start
from datasets import load_dataset
import… See the full description on the dataset page: https://huggingface.co/datasets/pulse-ai/PulseBench-Select.Pulse-Voltage-Response-Generation
