datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CT-RATE
The CT-RATE Team organizes the VLM3D Challenge
VLM3D 2026 (2nd Edition) → Challenge Finals at MICCAI 2026
VLM3D 2025 (1st Edition) → Challenge Finals at MICCAI 2025 • Workshop at ICCV 2025
The CT-RATE Team is developing the MR-RATE Dataset
A large-scale brain MRI dataset with paired radiology reports for training 3D vision-language models.
GitHub |
Dataset |
Metadata Dashboard
Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimhamamci/CT-RATE.GAMMA
GAMMA — Glaucoma grading from Multi-Modality imAges (Challenge dataset)
Image: Dataset Samples.
Short description
GAMMA is the first public multi-modality glaucoma grading dataset that pairs 2D color fundus photographs with 3D OCT volumes for each sample. It was released as part of the GAMMA challenge (OMIA8 / MICCAI 2021) to encourage algorithms that combine fundus and OCT information for automatic… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/GAMMA.CT_DeepLesion-MedSAM2
CT_DeepLesion-MedSAM2 Dataset
Authors
Jun Ma* 1,2,
Zongxin Yang* 3,
Sumin Kim2,4,5,
Bihui Chen2,4,5,
Mohammed Baharoon2,3,5,
Adibvafa Fallahpour2,4,5,
Reza Asakereh4,7,
Hongwei Lyu4,
Bo Wang† 1,2,4,5,6
* Equal contribution † Corresponding author
1AI Collaborative Centre, University Health Network, Toronto, Canada
2Vector… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/CT_DeepLesion-MedSAM2.CT_Lymph_Nodes
CT Lymph Nodes (Roth et al., NIH)
176 chest + abdomen CT scans with manually-traced voxel-wise mediastinal and
abdominal lymph node segmentations. The collection underpins the Roth 2014
detection benchmark and Seff 2015 segmentation benchmark, and remains a
widely-cited reference for thoracic-abdominal lymph node CADe work.
Dataset Details
Field
Value
Modality
CT
Body part
Mediastinum + Abdomen
Task
3D binary segmentation (foreground = lymph nodes)… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/CT_Lymph_Nodes.multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.CTVid-Bench
CTVid-Bench
CTVid-Bench is an open-source testing benchmark for clear text video restoration. This release packages the public evaluation media for three methods (GT, blur, downsample_x4) together with the latest QA v2 annotations.
Paper: ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement, accepted at ECCV 2026.
Project Page · arXiv:2608.28784 · Code
Release Scope
This folder is the Hugging Face staging… See the full description on the dataset page: https://huggingface.co/datasets/jinlong17/CTVid-Bench.dev_plantcad2_ft_long_ctxc4-10k-mini-tokenized-16-ctx-gelu-1l-testscode_x_glue_ct_code_to_text
Dataset Card for "code_x_glue_ct_code_to_text"
Dataset Summary
CodeXGLUE code-to-text dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Text/code-to-text
The dataset we use comes from CodeSearchNet and we filter the dataset as the following:
Remove examples that codes cannot be parsed into an abstract syntax tree.
Remove examples that #tokens of documents is < 3 or >256
Remove examples that documents contain special tokens (e.g. <img ...> or… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_ct_code_to_text.CT-RATE_Generated_Scans
Dataset Card for Synthetic Text-to-CT Scans - VLM3D Challenge
Dataset Details
Dataset Description
This dataset contains 1,000 synthetic 3D chest CT scans generated using the model introduced in
From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation (Molino et al., BMVC 2026).
The model was trained on the CT-RATE dataset, the largest publicly available collection of paired CT volumes and radiology reports.
It… See the full description on the dataset page: https://huggingface.co/datasets/dmolino/CT-RATE_Generated_Scans.alibaba-ctf
Alibaba CTF Benchmark
Alibaba CTF Benchmark is a CTF benchmark designed to measure the frontier of agent work on Capture The Flag security challenges. It consists of 87 high-quality tasks curated from the 2023–2026 AlibabaCTF (formerly AliyunCTF) competition series, covering five core categories: Web (25), Pwn (19), Misc (14), Reverse (16), and Crypto (13). During the curation process, LLM-based challenges were excluded due to their additional credential requirements and test… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/alibaba-ctf.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/DDR-dataset.CTSpine1KSpine-related diseases have high morbidity and cause a huge burden of social cost.
Spine imaging is an essential tool for noninvasively visualizing and assessing spinal
pathology. Segmenting vertebrae in computed tomography (CT) images is the basis of
quantitative medical image analysis for clinical diagnosis and surgery planning of
spine diseases. Current publicly available annotated datasets on spinal vertebrae are
small in size. Due to the lack of a large-scale annotated spine image dataset, the
mainstream deep learning-based segmentation methods, which are data-driven, are heavily
restricted. In this paper, we introduce a large-scale spine CT dataset, called CTSpine1K
curated from multiple sources for vertebra segmentation, which contains 1,005 CT volumes
with over 11,100 labeled vertebrae belonging to different spinal conditions. Based on
this dataset, we conduct several spinal vertebrae segmentation experiments to set the
first benchmark. We believe that this large-scale dataset will facilitate further
research in many spine-related image analysis tasks, including but not limited to
vertebrae segmentation, labeling, 3D spine reconstruction from biplanar radiographs,
image super-resolution, and enhancement.ctfhoard-corpusamos22-ct-dataset
AMOS22 CT Dataset
Dataset Description
This is the CT portion of the AMOS22 (A large-scale abdominal multi-organ benchmark for versatile medical image segmentation) dataset.
The AMOS22 dataset contains abdominal CT scans with dense segmentation annotations for 15 organs.
Dataset Structure
dict_keys(['train', 'valid']) splits:
train/
├── imagesTr/ # CT scan images in NIfTI format (.nii.gz)
└── labelsTr/ # Segmentation masks in NIfTI… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/amos22-ct-dataset.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.Pancreatic-CT-CBCT-SEG
Pancreatic-CT-CBCT-SEG
Breath-hold CT and cone-beam CT (CBCT) images with expert manual
organ-at-risk (OAR) segmentations from radiation treatments of locally
advanced pancreatic cancer at Memorial Sloan Kettering Cancer Center.
Dataset Details
Field
Value
Modality
CT (planning, breath-hold, contrast-enhanced) + CBCT (kV, deep-inspiration breath-hold)
Body part
Upper abdomen — gastrointestinal organs-at-risk
Task
3D multi-class segmentation (2… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Pancreatic-CT-CBCT-SEG.ctc-suite-eval
CTC suite eval ladders
The 22-task corpus-tracking-capacity suite: per-task context ladders from 2k to 1M tokens,
consumed by the ctc_suite task family on the prasann/ctc-suite branch of allenai/olmo-eval
(ctc_nq:r64k, suites ctc:figure / ctc:xlong / ctc:r128k / ...). One config per task, one
split per rung; each row is one unified-format example (documents + queries + answers + gold).
Public release note (2026-08-14). Gold answers are included — training on this data… See the full description on the dataset page: https://huggingface.co/datasets/PrasannSinghal/ctc-suite-eval.NanoJev-Data
NanoJev-Data — Unified game supervision and recorded evaluation
The complete data package for the current NanoJev model:
Maze, Snake, ViZDoom Basic and Predict Position. It includes the exact mixed
supervised-learning inputs, full expert episodes, frozen evaluation cohorts,
recorded comparisons, and the six-source hard Maze/Snake demonstration.
Training data
Split
Rows per hard/soft variant
Train
10,898
Dev
1,715
Calibration
1,709
Test
2,496
OOD… See the full description on the dataset page: https://huggingface.co/datasets/C-Tianyu/NanoJev-Data.apt-cti-reports
APT CTI Reports Dataset
A collection of 2,710 PDF reports on Advanced Persistent Threats (APT) and Cyber Threat Intelligence (CTI).
Structure
├── apt_groups/ (937 files) - Reports attributed to specific APT groups
└── other/ (1773 files) - Multi-attribution, unattributed, and general CTI reports
Filename Format
<GROUP/TYPE>__<YEAR>__<TITLE>.pdf
Examples:
APT28__2019__Fancy_Bear_Campaign.pdf
MULTI__2023__Global_Threat_Report.pdf… See the full description on the dataset page: https://huggingface.co/datasets/hackerman700000/apt-cti-reports.ct_scantrain_ctf_eefsozcu-news-2014(This dataset contains raw text, which are unlabeled.)
1,656 Turkish news articles from Sözcü Newspaper (http://www.sozcu.com.tr) between December 20, 2013, and March 11, 2014.
GitHub Repo: https://github.com/BilkentInformationRetrievalGroup/TUBITAK113E249/
If you would like to use any material in this repository, please cite this paper:
Toraman, C. and Can, F. (2017), Discovering story chains: A framework based on zigzagged search and news actors. Journal of the Association for… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/sozcu-news-2014.train_ctftb2-t4s-ctltest_ctfcti-bench
Dataset Card for CTIBench
A set of benchmark tasks designed to evaluate large language models (LLMs) on cyber threat intelligence (CTI) tasks.
Dataset Details
Dataset Description
CTIBench is a comprehensive suite of benchmark tasks and datasets designed to evaluate LLMs in the field of CTI.
Components:
CTI-MCQ: A knowledge evaluation dataset with multiple-choice questions to assess the LLMs' understanding of CTI standards, threats, detection strategies… See the full description on the dataset page: https://huggingface.co/datasets/AI4Sec/cti-bench.lodopab-ct-glimpse
LoDoPaB-CT subsets for GLIMPSE
Processed image subsets used by
GLIMPSE
(paper). Sinograms are not stored; they
are rendered on the fly by the GLIMPSE data pipeline (ODL / scikit-image).
Layout
Split
Contents
Format
train/
LoDoPaB-CT training slices
.npy float32 arrays
test/
LoDoPaB-CT test slices
.npy float32 arrays
ood/
out-of-distribution brain images
.jpg
import numpy as np # a .npy slice
img = np.load("train/0.npy") # (H, W) float32… See the full description on the dataset page: https://huggingface.co/datasets/AmirEhsan1995/lodopab-ct-glimpse.ctc-cell-cycle-hela
CTC Cell Cycle Dataset
Cell Tracking Challenge (CTC) live-cell microscopy with derived cell cycle state
labels for 3-class temporal classification.
What's actually hosted
The repo name says hela for historical reasons. Currently hosted: Fluo-N2DH-GOWT1
(GFP-tagged Oct4 in mouse embryonic stem cells), which is what the milestone baseline
trained on. HeLa data may be added later under a hela/ prefix.
Sequence
Frames
Used as
01/
92
training
02/
92
held-out… See the full description on the dataset page: https://huggingface.co/datasets/DnaRnaProteins/ctc-cell-cycle-hela.ctx
ctx
Find the cheapest AI coding setup that actually works on your repo.
CTX Fit analyzes your repository, tests promising AI coding configurations
against real tasks in it, and produces the winning configuration as a
reviewable change — in your working tree with --apply, or as a pull request
with --pr. It picks the cheapest setup that reliably works — reliability
is a requirement, not a tie-break — and if nothing beats what you already have,
it says so.
The winner is chosen… See the full description on the dataset page: https://huggingface.co/datasets/Stevesolun/ctx.
