datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GAMMA
GAMMA — Glaucoma grading from Multi-Modality imAges (Challenge dataset)
Image: Dataset Samples.
Short description
GAMMA is the first public multi-modality glaucoma grading dataset that pairs 2D color fundus photographs with 3D OCT volumes for each sample. It was released as part of the GAMMA challenge (OMIA8 / MICCAI 2021) to encourage algorithms that combine fundus and OCT information for automatic… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/GAMMA.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/DDR-dataset.PALM
🩺 PALM — Pathologic Myopia Fundus Image Dataset
Image: Dataset Samples.
📘 Overview
PALM (Pathologic Myopia) is a publicly available fundus image dataset developed for detecting pathologic myopia (PM) and analyzing associated retinal lesions and anatomical structures.
It was released for the Pathologic Myopia Challenge (PALM), hosted by the Chinese Academy of Sciences and Sun Yat-sen University, and… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/PALM.RFMID
🩺 RFMiD — Retinal Fundus Multi-Disease Image Dataset
Image: Dataset Samples.
The Retinal Fundus Multi-Disease Image Dataset (RFMiD) is designed for multi-disease detection and classification in retinal fundus photographs.It includes 3,200 high-quality color images with 46 labeled retinal disease conditions, curated by expert ophthalmologists from India.This dataset enables development of generalized deep learning… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/RFMID.EYEPACS
EyePACS — Diabetic Retinopathy Fundus Image Dataset
Image: EYEPACS Preprocess Samples.
📘 Overview
EyePACS (Eye Picture Archive Communication System) is a large-scale collection of retinal fundus images used for automated diabetic retinopathy (DR) detection.It formed the basis of the Kaggle Diabetic Retinopathy Detection challenge, enabling research into DR classification and screening.
The… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/EYEPACS.ctmatch_irCTMatch Information Retrieval Dataset
This is a dataset of processed clinical trials documents, somehwat of a duplication of that found in datasets/ir_datasets
except that these have been preprocessed with ctproc to clean and extract useful fields from the clinical trial documents.
Note: They are currently saved as text files because of the downstream task in ctmatch, though in the future they may be converted to .csv.
Each .txt file has exactly 374648 lines of corresponding data:… See the full description on the dataset page: https://huggingface.co/datasets/semaj83/ctmatch_ir.MR-RATE-nvseg-ctmr
MR-RATE: A Vision-Language Foundation Model and Dataset for Magnetic Resonance Imaging
This is the MR-RATE-nvseg-ctmr repository, part of the MR-RATE dataset release. It contains multi-label segmentations predicted with the NV-Segment-CTMR model. For full dataset details, native-space MRI volumes, radiology reports, metadata, and data splits, please refer to the MR-RATE repository. To explore, download, and work with the… See the full description on the dataset page: https://huggingface.co/datasets/Forithmus/MR-RATE-nvseg-ctmr.ctmatch_classificationCTMatch Classification Dataset
This is a combined set of 2 labelled datasets of:
topic (patient descriptions), doc (clinical trials documents - selected fields), and label ({0, 1, 2}) triples, in jsonl format.
(Somewhat of a duplication of some of the ir_dataset also available on HF.)
These have been processed using ctproc, and in this state can be used by various tokenizers for fine-tuning (see ctmatch for examples).
These 2 datasets contain no patient identifying information are openly… See the full description on the dataset page: https://huggingface.co/datasets/semaj83/ctmatch_classification.CTMDThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 598,
"total_tasks":1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zhaoraning/CTMD.ctm-affective
CTM Affective Benchmarks: MUStARD + UR-FUNNY
Preprocessed multimodal media and text splits for the two affective-reasoning
benchmarks used in the ctm-ai exp_affective experiments:
MUStARD — multimodal sarcasm detection from TV sitcom clips
(Friends, The Big Bang Theory, The Golden Girls, Sarcasmaholics).
UR-FUNNY — multimodal humor detection from TED talk clips.
This repo bundles the derived media (raw clips, muted video streams, and
extracted audio) alongside the JSON text… See the full description on the dataset page: https://huggingface.co/datasets/lwaekfjlk/ctm-affective.CTMDAAA2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 300,
"total_tasks":1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zhaoraning/CTMDAAA2.NV-Generate-CTMReval_CTMDAAA2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 292,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zhaoraning/eval_CTMDAAA2.CT-MRI-Atlas-Data
Subject TMF_005: High-Resolution Anatomical Brain Atlas
Overview
This dataset contains a 3D structural parcellation of a human brain derived from a clinical T1-weighted MRI scan. The processing pipeline transformed raw, non-standard clinical data into a 1mm isotropic isotropic volume and segmented it into 95+ distinct anatomical regions.
Data Provenance & Methodology
The data was processed using the FastSurfer pipeline (a deep-learning-based alternative to… See the full description on the dataset page: https://huggingface.co/datasets/triton7777/CT-MRI-Atlas-Data.meralion-asr-align-docker
meralion-asr-align (Docker image)
Prebuilt Docker image for the ASR-Align sidecar (CTC forced alignment + NeMo speaker diarization)
of the MERaLiON single-GPU ASR bundle. Shipped as a docker load-able tarball. No build required.
Use
huggingface-cli download MERaLiON-CTM/meralion-asr-align-docker \
meralion-asr-align.tar.gz --repo-type dataset --local-dir .
docker load -i meralion-asr-align.tar.gz # -> meralion/meralion-asr-align:latest
Bring it up via the… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON-CTM/meralion-asr-align-docker.ctms-multitask-sft-v6
CTMS Multi-task SFT — V6
A matched pair of corpora for a clinical-trial-management text-to-SQL agent, differing in
exactly one variable: whether generate_sql rows carry a <think> reasoning trace.
run_a (control)
run_b (traced)
total
23,049
23,049
train / val / test
18,698 / 2,172 / 2,179
18,698 / 2,172 / 2,179
traced train SQL rows
0
10,125 (81.0%)
gold SQL
identical, byte-for-byte
identical, byte-for-byte
Tasks
task
n… See the full description on the dataset page: https://huggingface.co/datasets/persistent-fm/ctms-multitask-sft-v6.CT_med
Contains:
TRIAL NAME
BRIEF
DRUG USED
DRUG CLASS
INDICATION
TARGET
THERAPY
LEAD SPONSOR
CT-MRI_metrics_validation_datasetCTMUmeralion-qwen-aligner-docker
meralion-qwen-aligner (Docker image)
Prebuilt Docker image for the Qwen3-ForcedAligner sidecar of the MERaLiON single-GPU ASR bundle
(CJK-routing force alignment). Shipped as a docker load-able tarball (the layers are large and slow
to push to Docker Hub from our region). No build required.
Use
huggingface-cli download MERaLiON-CTM/meralion-qwen-aligner-docker \
meralion-qwen-aligner.tar.gz --repo-type dataset --local-dir .
docker load -i… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON-CTM/meralion-qwen-aligner-docker.ctms-multitask-sft-v3
CTMS Multi-Task SFT — V3 (uppercase-Snowflake)
Supervised fine-tuning corpus for a Clinical Trial Management System (CTMS) analytics assistant, spanning
7 tasks over a 122-table CTMS schema. This is the V3 build: all SQL uses unquoted identifiers
that resolve against the uppercase-identifier Snowflake schema DUMMY_FORTREA_AI_MODEL.FORTREA_AI_MODEL_V3_CAP.
Data is fully synthetic (generated from a CTMS data generator). It contains no real patient,
investigator, or trial data.… See the full description on the dataset page: https://huggingface.co/datasets/persistent-fm/ctms-multitask-sft-v3.ctms-multitask-sft-v10-2ct_mod_processed_data_batch_1_evalct_mod_processed_data_batch_13_evalct_mod_processed_data_batch_16_evalct_mod_processed_data_batch_18_evalct_mod_processed_data_batch_21_evalct_mod_processed_data_batch_6_evalct_mod_processed_data_batch_7_evalct_mod_processed_data_batch_9_eval
