datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mimiciv_demoMIMIC-IVecg-qa-mimic-iv-ecg-250-500
Dataset Details
This is a randomly split five fold dataset of the MIMIC-IV-ECG varaint of ECG-QA stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-qa-mimic-iv-ecg-250-500 where 250 refers to the sampling frequency and 500 denotes 2 seconds.
The code to do the splitting is here.
The dataset contents are under the same license as the original ECG-QA's license.
Any questions or issues, please do not hesitate… See the full description on the dataset page: https://huggingface.co/datasets/willxxy/ecg-qa-mimic-iv-ecg-250-500.ecg-qa-mimic-iv-ecg-250-2500
Dataset Details
This is a randomly split five fold dataset of the MIMIC-IV-ECG varaint of ECG-QA stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-qa-mimic-iv-ecg-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
The dataset contents are under the same license as the original ECG-QA's license.
Any questions or issues, please do not… See the full description on the dataset page: https://huggingface.co/datasets/willxxy/ecg-qa-mimic-iv-ecg-250-2500.ecg-qa-mimic-iv-ecg-250-2500
Dataset Details
This is a randomly split five fold dataset of the MIMIC-IV-ECG varaint of ECG-QA stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-qa-mimic-iv-ecg-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
The dataset contents are under the same license as the original ECG-QA's license.
Any questions or issues, please do not… See the full description on the dataset page: https://huggingface.co/datasets/ELM-Research/ecg-qa-mimic-iv-ecg-250-2500.ecg-qa-mimic-iv-ecg-250-1250
Dataset Details
This is a randomly split five fold dataset of the MIMIC-IV-ECG varaint of ECG-QA stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-qa-mimic-iv-ecg-250-1250 where 250 refers to the sampling frequency and 1250 denotes 5 seconds.
The code to do the splitting is here.
The dataset contents are under the same license as the original ECG-QA's license.
Any questions or issues, please do not hesitate… See the full description on the dataset page: https://huggingface.co/datasets/willxxy/ecg-qa-mimic-iv-ecg-250-1250.mimic-iv-notesmimicivmimic-iv-echo-jepa-embeddings
MIMIC-IV-Echo V-JEPA & EchoJEPA Embeddings
Pre-computed video embeddings for all ~525K MIMIC-IV-Echo echocardiography clips using V-JEPA2 / V-JEPA2.1 and EchoJEPA fine-tuned variants.
Access requirement: You must have PhysioNet credentialed access and a signed DUA for MIMIC-IV-Echo before requesting access.
Models
Folder names map directly to source checkpoints. V-JEPA 2 and V-JEPA 2.1 are distinct — see the Description column. EchoJEPA models are all V-JEPA 2.1… See the full description on the dataset page: https://huggingface.co/datasets/MITCriticalData/mimic-iv-echo-jepa-embeddings.mimic-iv-noteecg-qa-mimic-iv-ecg-250-2500
Dataset Details
This is a randomly split five fold dataset of the MIMIC-IV-ECG varaint of ECG-QA stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-qa-mimic-iv-ecg-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
The dataset contents are under the same license as the original ECG-QA's license.
Any questions or issues, please do not… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/ecg-qa-mimic-iv-ecg-250-2500.MIMIC_IV_summarization_shortMIMIC_IV_Summarization_Samplemimiciv_meds_demoMIMIC_IV_lab_testsmimicIV-ginemimic_iv_word_countWord counts (excluding numerals and punctuation) of words from MIMIC-IV discharge summaries.
Citations
Johnson, A., Bulgarelli, L., Pollard, T., Gow, B., Moody, B., Horng, S., Celi, L. A., & Mark, R. (2024). MIMIC-IV (version 3.1). PhysioNet. RRID:SCR_007345.
MIMIC_IV_lab_test_individualMIMIC_IV_lab_test_individualmimic_iv_discharge_dutch_marianmt
Dataset Card for Mimic 4 Dutch Translation Marianmt
This dataset was created by: Laboratory for Computational Physiology, MIT Institute for Medical Engineering and Science,
in collaboration with the Beth Israel Deaconess Medical Center in Boston, Massachusetts.
The original data source: PhysioNet
Data description
Translation of MIMIC 4 using MariaNMT with vanilla settings.
Note: this version contains spurious repetitions that will need to be filtered using e.g. regular… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/mimic_iv_discharge_dutch_marianmt.mimic_iv_radio_dutch_marianmt
Dataset Card for Mimic 4 Dutch Translation Marianmt
This dataset was created by: Laboratory for Computational Physiology, MIT Institute for Medical Engineering and Science,
in collaboration with the Beth Israel Deaconess Medical Center in Boston, Massachusetts.
The original data source: PhysioNet
Data description
Translation of MIMIC 4 using MariaNMT with vanilla settings.
Note: this version contains spurious repetitions that will need to be filtered using e.g. regular… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/mimic_iv_radio_dutch_marianmt.
