datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CRCICLR
CRCICLR / HSC-TTA EEG research artifacts
This public dataset repository contains code, experiment outputs, logs, configurations,
model artifacts, data manifests, and derived/processed features from the CRCICLR HSC-TTA
EEG research project.
Original raw EEG files are intentionally excluded. Credentials, Git history, Conda
environments, caches, private keys, and temporary files are also excluded.
The server-side source files were retained; this publication is a copy-only… See the full description on the dataset page: https://huggingface.co/datasets/gifoe/CRCICLR.CRCD
Comprehensive Robotic Cholecystectomy Dataset (CRCD)
The Comprehensive Robotic Cholecystectomy Dataset (CRCD) is a large-scale, multimodal dataset for robot-assisted surgery (RAS) research.It provides synchronized endoscopic videos, da Vinci surgical robot kinematics, and pedal usage signals, making it one of the most comprehensive open datasets for studying robotic cholecystectomy procedures.
CRCD supports research in:
Medical robotics and surgical automation
Computer vision… See the full description on the dataset page: https://huggingface.co/datasets/SITL-Eng/CRCD.NCT-CRC-HE
100,000 histological images of human colorectal cancer and healthy tissue
Data Description "NCT-CRC-HE-100K"
This is a set of 100,000 non-overlapping image patches from hematoxylin & eosin (H&E) stained histological images of human colorectal cancer (CRC) and normal tissue.
All images are 224x224 pixels (px) at 0.5 microns per pixel (MPP). All images are color-normalized using Macenko's method (http://ieeexplore.ieee.org/abstract/document/5193250/, DOI… See the full description on the dataset page: https://huggingface.co/datasets/1aurent/NCT-CRC-HE.CRC_comparativeGAEA-Train GAEA: A Geolocation Aware Conversational Model [WACV 2026🔥]
Summary
Image geolocalization, in which an AI model traditionally predicts the precise GPS coordinates of an image, is a challenging task with many downstream applications. However, the user cannot utilize the model to further their knowledge beyond the GPS coordinates; the model lacks an understanding of the location and the conversational ability to communicate with the user. In recent days, with the tremendous progress of… See the full description on the dataset page: https://huggingface.co/datasets/ucf-crcv/GAEA-Train.crc1-3d-histology-modulecrcis-quranic-eval-leaderboard-results_details_PartAI__Dorna-Llama3-8B-Instruct_private
Dataset Card for Evaluation run of PartAI/Dorna-Llama3-8B-Instruct
Dataset automatically created during the evaluation run of model PartAI/Dorna-Llama3-8B-Instruct.
The dataset is composed of 8 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/sadra-barikbin/crcis-quranic-eval-leaderboard-results_details_PartAI__Dorna-Llama3-8B-Instruct_private.NCT-CRC-HE
100,000 histological images of human colorectal cancer and healthy tissue
Data Description "NCT-CRC-HE-100K"
This is a set of 100,000 non-overlapping image patches from hematoxylin & eosin (H&E) stained histological images of human colorectal cancer (CRC) and normal tissue.
All images are 224x224 pixels (px) at 0.5 microns per pixel (MPP). All images are color-normalized using Macenko's method (http://ieeexplore.ieee.org/abstract/document/5193250/, DOI… See the full description on the dataset page: https://huggingface.co/datasets/yh123yh/NCT-CRC-HE.crcis-quranic-eval-leaderboard-results_details_mistralai__Mistral-7B-v0.1_private
Dataset Card for Evaluation run of mistralai/Mistral-7B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-v0.1.
The dataset is composed of 6 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 16 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/sadra-barikbin/crcis-quranic-eval-leaderboard-results_details_mistralai__Mistral-7B-v0.1_private.NCT-CRC-HE
100,000 histological images of human colorectal cancer and healthy tissue
Data Description "NCT-CRC-HE-100K"
This is a set of 100,000 non-overlapping image patches from hematoxylin & eosin (H&E) stained histological images of human colorectal cancer (CRC) and normal tissue.
All images are 224x224 pixels (px) at 0.5 microns per pixel (MPP). All images are color-normalized using Macenko's method (http://ieeexplore.ieee.org/abstract/document/5193250/, DOI… See the full description on the dataset page: https://huggingface.co/datasets/Varsha-Y12/NCT-CRC-HE.crc1-3d-histology-module-v2nct-crc-he
Dataset Card for NCT-CRC-HE
Dataset Summary
The NCT-CRC-HE dataset consists of images of human tissue slides, some of which contain cancer.
Data Splits
The dataset contains tissues from different parts of the body. Examples from each of the 9 classes can be seen below
Initial Data Collection and Normalization
NCT biobank (National Center for Tumor Diseases) and the UMM pathology archive (University Medical Center Mannheim). Images were… See the full description on the dataset page: https://huggingface.co/datasets/owkin/nct-crc-he.Tahoe-x1-CRC-embeddings
Tahoe-x1 3B — CRC slice
Cell embeddings from tahoebio/Tahoe-x1-embeddings
(Tx1-3B on Tahoe-100M), restricted to CRC cell lines.
This is a filter of the published embeddings, not a new model run.
Source license: Apache-2.0.
Cells: 21,068,133
line
CVCL
n cells
SW480
CVCL_0546
6,040,371
LoVo
CVCL_0399
3,013,246
RKO
CVCL_0504
2,182,314
HT-29
CVCL_0320
2,171,036
SW1417
CVCL_1717
1,921,742
LS 180
CVCL_0397
1,828,299
HCT15
CVCL_0292
1,500,121
COLO 205
CVCL_0218… See the full description on the dataset page: https://huggingface.co/datasets/dn-gh/Tahoe-x1-CRC-embeddings.BBQ-V
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
⚠️ Content warning: This dataset contains contexts and questions that surface
harmful social stereotypes. It is intended solely for measuring and mitigating bias
in AI systems.
Summary
Stereotype biases in Large Multimodal Models (LMMs) perpetuate harmful societal prejudices, undermining the fairness and equity of AI applications. As LMMs grow increasingly influential, addressing and… See the full description on the dataset page: https://huggingface.co/datasets/ucf-crcv/BBQ-V.HEPROBench-CRC-CODEX-review-demo
HEPROBench CRC-CODEX reviewer demo
This is a small real-data software-verification subset for HEPROBench. It
contains paired, registered H&E and CRC-CODEX patches derived from Schürch et
al., Coordinated cellular neighborhoods orchestrate antitumoral immunity at
the colorectal cancer invasive front, created by Christian Schürch, Mendeley Data
10.17632/mpjzbtfgfr.1, licensed under
CC BY 4.0.
The bundle has 2 patches from each of 2
anonymized FOVs per split (4 train,
4 validation… See the full description on the dataset page: https://huggingface.co/datasets/u3011706/HEPROBench-CRC-CODEX-review-demo.GAEA-Bench GAEA: A Geolocation Aware Conversational Assistant [WACV 2026🔥]
Summary
Image geolocalization, in which an AI model traditionally predicts the precise GPS coordinates of an image, is a challenging task with many downstream applications. However, the user cannot utilize the model to further their knowledge beyond the GPS coordinates; the model lacks an understanding of the location and the conversational ability to communicate with the user. In recent days, with the tremendous progress of… See the full description on the dataset page: https://huggingface.co/datasets/ucf-crcv/GAEA-Bench.crcis-quranic-eval-leaderboard-results_details_google__gemma-2-27b-it_private
Dataset Card for Evaluation run of google/gemma-2-27b-it
Dataset automatically created during the evaluation run of model google/gemma-2-27b-it.
The dataset is composed of 7 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/sadra-barikbin/crcis-quranic-eval-leaderboard-results_details_google__gemma-2-27b-it_private.CRCD-dVRK-LeRobot
CRCD-dVRK-LeRobot
18 episodes of robot-assisted cholecystectomy on pig liver using the da Vinci Research Kit (dVRK). LeRobot v3.0 format.
Converted from the CRCD dataset (University of Verona / Bologna).
Episodes
18 (7 surgeons, 1-3 attempts each)
Frames
377,950
FPS
30
Robot
dVRK
Task
Cholecystectomy (ex vivo, pig liver)
Camera
Stereo endoscope, 848x480
State
16D — PSM1/PSM2 cartesian pose (xyz + quaternion) + jaw
Action
16D — next-timestep absolute… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz045/CRCD-dVRK-LeRobot.TF-CoVR From Play to Replay: Composed Video Retrieval for
Temporally Fine-Grained Videos
Accepted in NeurIPS 2025
Animesh Gupta1 |
Jay Parmar1 |
Ishan Rajendrakumar Dave2 |
Mubarak Shah1
1University of Central Florida 2Adobe
Abstract
Composed Video Retrieval (CoVR) retrieves a target video given a query video and a modification text describing the intended change. Existing CoVR benchmarks emphasize appearance shifts or coarse event changes and… See the full description on the dataset page: https://huggingface.co/datasets/ucf-crcv/TF-CoVR.gtcapiotsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 100,
"total_frames": 43546,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/crc-amd-hackathon-2025/gtcapiots.CRCD-dVRK-LeRobot
CRCD-dVRK-LeRobot
18 episodes of robot-assisted cholecystectomy on pig liver using the da Vinci Research Kit (dVRK). LeRobot v3.0 format.
Converted from the CRCD dataset (University of Verona / Bologna).
Episodes
18 (7 surgeons, 1-3 attempts each)
Frames
377,950
FPS
30
Robot
dVRK
Task
Cholecystectomy (ex vivo, pig liver)
Camera
Stereo endoscope, 848x480
State
16D — PSM1/PSM2 cartesian pose (xyz + quaternion) + jaw
Action
16D — next-timestep absolute… See the full description on the dataset page: https://huggingface.co/datasets/morozovdd/CRCD-dVRK-LeRobot.datasetimage
crc1-3d-histology-siteCRCD-sentiment-balanced-3class
CRCD Balanced Sentiment Dataset
A cleaned and balanced English dataset for three-class sentiment
classification of customer and product reviews.
Dataset contents
Each record contains two fields:
text — the review text
label — the numerical sentiment label
Label mapping
Label
Sentiment
0
negative
1
neutral
2
positive
Class distribution
Split
Total
Negative
Neutral
Positive
train
5,551
1,851
1,850
1,850… See the full description on the dataset page: https://huggingface.co/datasets/SergeiM89/CRCD-sentiment-balanced-3class.miniMTI-CRC-example
miniMTI-CRC Example Data
Example single-cell imaging data for testing miniMTI, a molecularly anchored virtual staining framework for multiplex tissue imaging panel reduction.
Paper: bioRxiv 2026.01.21.700911Code: GitHubModel: changlab/miniMTI-CRC
Dataset Description
10,000 single-cell image patches randomly sampled (seed=42) from CRC-Orion sample CRC04 (colorectal cancer tissue WSI, RareCyte Orion platform).
File
example_CRC04_10k.h5 — HDF5 file (~178… See the full description on the dataset page: https://huggingface.co/datasets/changlab/miniMTI-CRC-example.ImplicitQA VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
Sirnam Swetha |
Rohit Gupta |
Parth Parag Kulkarni |
David G Shatwell |
Jeffrey A Chan Santiago |
Nyle Siddiqui |
Joseph Fioresi |
Mubarak Shah
University of Central Florida
VRRQA Dataset
The VRRQA dataset was introduced in the paper VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues.
Project page:… See the full description on the dataset page: https://huggingface.co/datasets/ucf-crcv/ImplicitQA.CRC100kcrc1maio-assetsgrab-camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 74,
"total_frames": 40044,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:74"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/crc-amd-hackathon-2025/grab-cam.crcis-quranic-eval-leaderboard-results_details_openai__gpt-4o-mini_private
Dataset Card for Evaluation run of OpenAI/gpt-4o-mini
Dataset automatically created during the evaluation run of model OpenAI/gpt-4o-mini.
The dataset is composed of 7 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/sadra-barikbin/crcis-quranic-eval-leaderboard-results_details_openai__gpt-4o-mini_private.
