datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CRCD
Comprehensive Robotic Cholecystectomy Dataset (CRCD)
The Comprehensive Robotic Cholecystectomy Dataset (CRCD) is a large-scale, multimodal dataset for robot-assisted surgery (RAS) research.It provides synchronized endoscopic videos, da Vinci surgical robot kinematics, and pedal usage signals, making it one of the most comprehensive open datasets for studying robotic cholecystectomy procedures.
CRCD supports research in:
Medical robotics and surgical automation
Computer vision… See the full description on the dataset page: https://huggingface.co/datasets/SITL-Eng/CRCD.GAEA-Train GAEA: A Geolocation Aware Conversational Model [WACV 2026🔥]
Summary
Image geolocalization, in which an AI model traditionally predicts the precise GPS coordinates of an image, is a challenging task with many downstream applications. However, the user cannot utilize the model to further their knowledge beyond the GPS coordinates; the model lacks an understanding of the location and the conversational ability to communicate with the user. In recent days, with the tremendous progress of… See the full description on the dataset page: https://huggingface.co/datasets/ucf-crcv/GAEA-Train.crcis-quranic-eval-leaderboard-results_details_PartAI__Dorna-Llama3-8B-Instruct_private
Dataset Card for Evaluation run of PartAI/Dorna-Llama3-8B-Instruct
Dataset automatically created during the evaluation run of model PartAI/Dorna-Llama3-8B-Instruct.
The dataset is composed of 8 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/sadra-barikbin/crcis-quranic-eval-leaderboard-results_details_PartAI__Dorna-Llama3-8B-Instruct_private.crcis-quranic-eval-leaderboard-results_details_mistralai__Mistral-7B-v0.1_private
Dataset Card for Evaluation run of mistralai/Mistral-7B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-v0.1.
The dataset is composed of 6 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 16 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/sadra-barikbin/crcis-quranic-eval-leaderboard-results_details_mistralai__Mistral-7B-v0.1_private.Tahoe-x1-CRC-embeddings
Tahoe-x1 3B — CRC slice
Cell embeddings from tahoebio/Tahoe-x1-embeddings
(Tx1-3B on Tahoe-100M), restricted to CRC cell lines.
This is a filter of the published embeddings, not a new model run.
Source license: Apache-2.0.
Cells: 21,068,133
line
CVCL
n cells
SW480
CVCL_0546
6,040,371
LoVo
CVCL_0399
3,013,246
RKO
CVCL_0504
2,182,314
HT-29
CVCL_0320
2,171,036
SW1417
CVCL_1717
1,921,742
LS 180
CVCL_0397
1,828,299
HCT15
CVCL_0292
1,500,121
COLO 205
CVCL_0218… See the full description on the dataset page: https://huggingface.co/datasets/dn-gh/Tahoe-x1-CRC-embeddings.BBQ-V
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
⚠️ Content warning: This dataset contains contexts and questions that surface
harmful social stereotypes. It is intended solely for measuring and mitigating bias
in AI systems.
Summary
Stereotype biases in Large Multimodal Models (LMMs) perpetuate harmful societal prejudices, undermining the fairness and equity of AI applications. As LMMs grow increasingly influential, addressing and… See the full description on the dataset page: https://huggingface.co/datasets/ucf-crcv/BBQ-V.GAEA-Bench GAEA: A Geolocation Aware Conversational Assistant [WACV 2026🔥]
Summary
Image geolocalization, in which an AI model traditionally predicts the precise GPS coordinates of an image, is a challenging task with many downstream applications. However, the user cannot utilize the model to further their knowledge beyond the GPS coordinates; the model lacks an understanding of the location and the conversational ability to communicate with the user. In recent days, with the tremendous progress of… See the full description on the dataset page: https://huggingface.co/datasets/ucf-crcv/GAEA-Bench.crcis-quranic-eval-leaderboard-results_details_google__gemma-2-27b-it_private
Dataset Card for Evaluation run of google/gemma-2-27b-it
Dataset automatically created during the evaluation run of model google/gemma-2-27b-it.
The dataset is composed of 7 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/sadra-barikbin/crcis-quranic-eval-leaderboard-results_details_google__gemma-2-27b-it_private.TF-CoVR From Play to Replay: Composed Video Retrieval for
Temporally Fine-Grained Videos
Accepted in NeurIPS 2025
Animesh Gupta1 |
Jay Parmar1 |
Ishan Rajendrakumar Dave2 |
Mubarak Shah1
1University of Central Florida 2Adobe
Abstract
Composed Video Retrieval (CoVR) retrieves a target video given a query video and a modification text describing the intended change. Existing CoVR benchmarks emphasize appearance shifts or coarse event changes and… See the full description on the dataset page: https://huggingface.co/datasets/ucf-crcv/TF-CoVR.CRCD-sentiment-balanced-3class
CRCD Balanced Sentiment Dataset
A cleaned and balanced English dataset for three-class sentiment
classification of customer and product reviews.
Dataset contents
Each record contains two fields:
text — the review text
label — the numerical sentiment label
Label mapping
Label
Sentiment
0
negative
1
neutral
2
positive
Class distribution
Split
Total
Negative
Neutral
Positive
train
5,551
1,851
1,850
1,850… See the full description on the dataset page: https://huggingface.co/datasets/SergeiM89/CRCD-sentiment-balanced-3class.miniMTI-CRC-example
miniMTI-CRC Example Data
Example single-cell imaging data for testing miniMTI, a molecularly anchored virtual staining framework for multiplex tissue imaging panel reduction.
Paper: bioRxiv 2026.01.21.700911Code: GitHubModel: changlab/miniMTI-CRC
Dataset Description
10,000 single-cell image patches randomly sampled (seed=42) from CRC-Orion sample CRC04 (colorectal cancer tissue WSI, RareCyte Orion platform).
File
example_CRC04_10k.h5 — HDF5 file (~178… See the full description on the dataset page: https://huggingface.co/datasets/changlab/miniMTI-CRC-example.ImplicitQA VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
Sirnam Swetha |
Rohit Gupta |
Parth Parag Kulkarni |
David G Shatwell |
Jeffrey A Chan Santiago |
Nyle Siddiqui |
Joseph Fioresi |
Mubarak Shah
University of Central Florida
VRRQA Dataset
The VRRQA dataset was introduced in the paper VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues.
Project page:… See the full description on the dataset page: https://huggingface.co/datasets/ucf-crcv/ImplicitQA.CRC100kcrcis-quranic-eval-leaderboard-results_details_openai__gpt-4o-mini_private
Dataset Card for Evaluation run of OpenAI/gpt-4o-mini
Dataset automatically created during the evaluation run of model OpenAI/gpt-4o-mini.
The dataset is composed of 7 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/sadra-barikbin/crcis-quranic-eval-leaderboard-results_details_openai__gpt-4o-mini_private.teccrcis-quranic-eval-leaderboard-results_details_CRCIS__T2-Dorna-6000Ada-QA_private
Dataset Card for Evaluation run of CRCIS/T2-Dorna-6000Ada-QA
Dataset automatically created during the evaluation run of model CRCIS/T2-Dorna-6000Ada-QA.
The dataset is composed of 4 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/sadra-barikbin/crcis-quranic-eval-leaderboard-results_details_CRCIS__T2-Dorna-6000Ada-QA_private.CrCoNi_Cao_2022
Cite this dataset Cao, Y., Sheriff, K., and Freitas, R. CrCoNi Cao 2022. ColabFit, 2024. https://doi.org/10.60732/76208b62
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_z2mkok0egrm8_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.
https://materials.colabfit.org… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/CrCoNi_Cao_2022.emo_isCRC_TEST_CODESPACE
Unified Seismic Soil Liquefaction Case-History Dataset
I put this together because the liquefaction case histories I needed for my own work
were scattered across about a dozen spreadsheets, every one of them with different
units, column names and label conventions. Merging them by hand every time got old,
so I cleaned everything once into a single CSV and figured other people might as well
use it too.
It's 1,830 field case histories from 9 public sources, all re-expressed under… See the full description on the dataset page: https://huggingface.co/datasets/guanwencan/CRC_TEST_CODESPACE.china-al-crc-steel-price-data-july-2026
China AL and CRC Historical Steel Price Data – July 2026
This dataset provides historical price range data for AL (Aluminum Coil) and CRC (Cold Rolled Coil) products in mainland China during July 2026.
The dataset includes daily low and high price ranges, currency, unit, market information, and source references.
The data is compiled and maintained by iPPGI for steel market research, historical price comparison, trend analysis, and data visualization.
Data Source… See the full description on the dataset page: https://huggingface.co/datasets/ippgidata/china-al-crc-steel-price-data-july-2026.crc_image_dataset
Dataset Card for "crc_image_dataset"
More Information needed
ai-product-crc-trainingcrcs-provenance-trailcrdflowercrc-devcrc-ai-csvcolon-3d-crc1stimutestcrc-ai-data
