datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ecg-logECGInstruct
ECGInstruct
Dataset for paper "Teach Multimodal LLMs to Comprehend Electrocardiographic Images".
🌐 Project Page: https://aimedlab.github.io/PULSE/
📄 Paper: https://arxiv.org/abs/2410.19008
🧑💻 Code: https://github.com/AIMedLab/PULSE
🤗 Model: https://huggingface.co/PULSE-ECG/PULSE-7B
⚖️ ECGBench: https://huggingface.co/datasets/PULSE-ECG/ECGBench
Introduction
ECGInstruct is a comprehensive and large-scale instruction-tuning dataset designed for ECG image… See the full description on the dataset page: https://huggingface.co/datasets/PULSE-ECG/ECGInstruct.pulsefeed-x402-security
PulseFeed — x402 Agent-Payment Security & Trust (open data)
Independent, daily-updated trust & safety data for the x402 agent-payment economy (HTTP 402 + stablecoins on Base) and the MCP server ecosystem — by PulseFeed.
AI agents increasingly pay for APIs autonomously over x402 and connect to MCP servers that can run code on install. But 20% of listed x402 endpoints are dead or invalid, and "live" is not the same claim as "payable": of 25709 endpoints that return a valid 402… See the full description on the dataset page: https://huggingface.co/datasets/Nikolife/pulsefeed-x402-security.PulseLM
PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning
Usage
from datasets import load_dataset, get_dataset_config_names, concatenate_datasets # datasets==4.5.0
dataset_names = get_dataset_config_names("Manhph2211/PulseLM")
print(f"Available datasets: {dataset_names}")
train_splits = [
load_dataset("Manhph2211/PulseLM", name, split="train").select_columns(["signal", "text", "qa"])
for name in dataset_names
]
combined =… See the full description on the dataset page: https://huggingface.co/datasets/Manhph2211/PulseLM.pulseedge-trading-ai-v5
PulseEdge Trading AI v5 — categorized lake
Private snapshot of the Hetzner storage box, split by data type.
v4 put packed tables at the repo root and everything else under a flat archives/ dump.
v5 keeps the same research lake but files live in typed folders.
Research use only. Not investment advice. Equities/ETFs only. No crypto.
Categories
Folder
What is in it
ohlcv/packed/
Combined Yahoo bars (1d/1h/15m/5m/1m) plus train/holdout splits
ohlcv/yahoo/… See the full description on the dataset page: https://huggingface.co/datasets/NickTheCoolst/pulseedge-trading-ai-v5.pulseedge-trading-ai-v2
PulseEdge Trading AI v2
Broader follow-up to NickTheCoolst/pulseedge-trading-ai (v1).
Equities/ETFs only. No crypto. No raw ITCH. No SEC archives.
Research use only. Not investment advice. Survivorship bias in the liquidity universe.
Use purged walk-forward splits. Label columns leak the future by construction.
What is new vs v1
v1
v2
Liquidity bar
$250k, 2y, $2
$80k, 1y, $1
Daily features (filtered)
24.1M / 6,896
25.6M / 7,505
Daily features (full… See the full description on the dataset page: https://huggingface.co/datasets/NickTheCoolst/pulseedge-trading-ai-v2.x402-market-pulse
x402 Market Pulse
A small, growing time series of how much USDC actually settles through x402 pay-per-call APIs on Base,
measured from on-chain Transfer logs rather than from provider-reported call counts. One row per
measurement; every row is a trailing 24-hour window ending at measured_at_utc.
Status: bootstrapping. 71 rows spanning 2026-09-22T13:34Z to 2026-09-25T15:00Z —
less than one independent day of data so far. Rows are kept to at most one per clock hour and a new one… See the full description on the dataset page: https://huggingface.co/datasets/nimapro1381/x402-market-pulse.ECGBench
ECGBench
Benchmark Dataset for paper "Teach Multimodal LLMs to Comprehend Electrocardiographic Images".
🌐 Project Page: https://aimedlab.github.io/PULSE/
📄 Paper: https://arxiv.org/abs/2410.19008
🧑💻 Code: https://github.com/AIMedLab/PULSE
🤗 Model: https://huggingface.co/PULSE-ECG/PULSE-7B
👩⚕️ ECGInstruct: https://huggingface.co/datasets/PULSE-ECG/ECGInstruct
Introduction
We introduce ECGBench, a comprehensive benchmark designed to evaluate ECG image… See the full description on the dataset page: https://huggingface.co/datasets/PULSE-ECG/ECGBench.ecg-instruct-pulse-250-500
Dataset Details
This is a randomly split five fold dataset of the ECG-Instruct Pulse dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-instruct-pulse-250-500 where 250 refers to the sampling frequency and 500 denotes 2 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.
llm-pulse-dataPULSE
PULSE
A Synchronized Five-Modality Dataset for Multi-Modal Daily Activity Understanding
MoCap · EMG · Eye Tracking · IMU · Fingertip Pressure — all hardware-synced at 100 Hz
NeurIPS 2026 Evaluations & Datasets · under double-blind review · CC BY-NC 4.0
At a glance
40
9
5
7,789
volunteers
scenarios (S1–S8 + S9 motion primitives)
modalities @ 100 Hz
dense action segments
337 total recordings(304 task + 33 S9)
~9.7 h total (S1–S8: ~7.0 h)
<10 ms… See the full description on the dataset page: https://huggingface.co/datasets/velvet-pine-22/PULSE.ecg-instruct-pulse-250-1250
Dataset Details
This is a randomly split five fold dataset of the ECG-Instruct Pulse dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-instruct-pulse-250-1250 where 250 refers to the sampling frequency and 1250 denotes 5 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.
pulse-sofroniew-emotion-concept-texts
Pulse Geometry: Sofroniew-Style Implicit Emotion Corpus
A contrastive corpus of 8,550 short stories (171 emotions × 50 topics) that convey
a target emotion implicitly — through behavior, sensation, dialogue, internal
thought, or environmental description, but never by naming the emotion. Each story
is scored on a four-axis rubric by Claude Sonnet.
The corpus was built as the substrate for a geometry replication: probing whether
an emotion-vector layout analogous to Sofroniew et al.… See the full description on the dataset page: https://huggingface.co/datasets/jmccardle/pulse-sofroniew-emotion-concept-texts.PulseBench-Tab
PulseBench-Tab
A frontier multilingual benchmark for table extraction from document images.
PulseBench-Tab contains 1,820 human-annotated tables across 9 languages (English 589, Chinese 213, Spanish 176, Russian 170, French 165, Japanese 164, Arabic 146, German 113, Korean 84) and 4 scripts (Latin, CJK, Arabic, Cyrillic), sourced from 380 unique documents including financial filings, government reports, corporate disclosures, and regulatory filings. Each sample is a table image… See the full description on the dataset page: https://huggingface.co/datasets/pulse-ai/PulseBench-Tab.fed-pulse-embedding-caches
yusufizzetmurat/fed-pulse-embedding-caches
Artefact for the fed-pulse FOMC text analytics project.
Source code: https://github.com/yusufizzetmurat/fed-pulse
Live demo: https://fedpulse.yusufizzetmurat.com
License: mit
Training corpus
Per-encoder embedding caches keyed on (encoder alias, encoder revision). Each parquet carries record_id, doc_id, event_date, chunk_index, chunk_preview, and embedding.
Training command
python… See the full description on the dataset page: https://huggingface.co/datasets/yusufizzetmurat/fed-pulse-embedding-caches.ultimate-life-ds15-volcanic-pulse
Ultimate Life — Volcanic Pulse
A public, U.S.-first, source-derived visual-grammar reference pack for MiniMax H3-compatible video study.
What is inside
clips/ — 20 silent H.264 MP4 clips and 20 paired .txt captions.
manifest.json — exact source timestamp, output hash, H3 contract, and clip-level restriction.
source.json — source URL, rights basis, preservation checksum, and probe data when available.
candidate-beats.json — selection ledger used to make the… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/ultimate-life-ds15-volcanic-pulse.pulse-ecg-instruct-subsetopen-pulse-hackathon-data-analysis
LauzHack Projects Dataset
Dataset Summary
This dataset contains comprehensive information about projects submitted to
LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project
includes details about the project title, description, team members, awards, and
categories.
LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique
Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and
hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.londons-pulse-moh
London's Pulse: Medical Officer of Health reports (page images + OCR text)
Page-level scans of the Wellcome Collection London's Pulse
Medical Officer of Health (MOH) reports (1848–1972), paired with OCR text, per-report
licence, and full provenance. Built for OCR / VLM / document-understanding work on real
historical public-health records — dense statistical tables, mixed layouts, century-old print.
Configs
config
rows
what
default
391,964 pages / 4,886… See the full description on the dataset page: https://huggingface.co/datasets/biglam/londons-pulse-moh.ecg-instruct-pulse-250-2500
Dataset Details
This is a randomly split five fold dataset of the ECG-Instruct Pulse dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-instruct-pulse-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.
worldcup-pulse-data
WorldCup Pulse Data
This Hugging Face Dataset repository is the single source of truth for WorldCup Pulse Lakehouse data. The GitHub code repo never commits generated data; GitHub Actions uploads this lakehouse output to this dataset repo.
Lakehouse layout
bronze/*.parquet # normalized raw API snapshots
silver/*.parquet # cleaned canonical entities and events
gold/*.parquet # dashboard-ready marts
state/last_run.json # incremental state
logs/*.csv… See the full description on the dataset page: https://huggingface.co/datasets/n2d/worldcup-pulse-data.PulseLM
PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning
Hung Manh Pham*
Jinyang Wu*
Xiao Ma
Yiming Zhang
Yixin Xu
Aaqib Saeed
Bin Zhu†
Zhou Pan†
Dong Ma†
* Equal contribution † Corresponding authors
Introduction
PulseLM is a multimodal framework that integrates PPG (Photoplethysmography) signal encoders with large language models for physiological signal understanding research. The project includes a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/Ronilos/PulseLM.ecg-bench-pulse-250-1250
Dataset Details
This is a randomly split five fold dataset of the ECG-Bench Pulse dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-bench-pulse-250-1250 where 250 refers to the sampling frequency and 1250 denotes 5 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.
ecg-bench-pulse-250-500
Dataset Details
This is a randomly split five fold dataset of the ECG-Bench Pulse dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-bench-pulse-250-500 where 250 refers to the sampling frequency and 500 denotes 2 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.
full-dubai-pulseecg-bench-pulse-250-2500
Dataset Details
This is a randomly split five fold dataset of the ECG-Bench Pulse dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-bench-pulse-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.
PulseMind
MediScope: A Large-Scale Multimodal Medical Dataset
MediScope is a large-scale multimodal medical dataset designed for medical vision-language understanding and medical QA. It includes structured JSON annotations paired with medical images.
Release Notes
This release provides a curated subset of approximately 1,000 cases (JSON + images).The full dataset is larger and will be gradually released in future updates.
Benchmarks
We provide the following benchmark… See the full description on the dataset page: https://huggingface.co/datasets/AQ-MedAI/PulseMind.fed-pulse-training-package
yusufizzetmurat/fed-pulse-training-package
Artefact for the fed-pulse FOMC text analytics project.
Source code: https://github.com/yusufizzetmurat/fed-pulse
Live demo: https://fedpulse.yusufizzetmurat.com
License: mit
Training corpus
Canonical FOMC event dataset: events.parquet, splits, fold manifest, rates_panel.parquet, linguistic_features.parquet, mp_surprises.parquet, macro_state.parquet, registry_normalized.jsonl, and quality reports.
Training… See the full description on the dataset page: https://huggingface.co/datasets/yusufizzetmurat/fed-pulse-training-package.Nexus-Pulse-WA2-Track1
Nexus-Pulse
WorldArena 2.0 Track 1 package submitted by Lei Yang, Tsinghua University.
Model name: Nexus-Pulse
Version: v2-rank03
Inference seed: 42
Contact: yanglei20@mails.tsinghua.edu.cn
Number of videos: 1000
Resolution: 640 x 480
Frames per video: 121
Frame rate: 24 FPS
This package is a stochastic inference variant of the frozen W0 checkpoint.
It is one of multiple transparently disclosed inference-seed variants of the
same base checkpoint, not an independently trained… See the full description on the dataset page: https://huggingface.co/datasets/yanglei18/Nexus-Pulse-WA2-Track1.120-physicsdojo137-pulseox-test4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 20,
"total_frames": 19763,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/120-physicsdojo137-pulseox-test4.
