CTC
Datasets
All datasets matching “CTC”ctc-suite-eval
CTC suite eval ladders
The 22-task corpus-tracking-capacity suite: per-task context ladders from 2k to 1M tokens,
consumed by the ctc_suite task family on the prasann/ctc-suite branch of allenai/olmo-eval
(ctc_nq:r64k, suites ctc:figure / ctc:xlong / ctc:r128k / ...). One config per task, one
split per rung; each row is one unified-format example (documents + queries + answers + gold).
Public release note (2026-08-14). Gold answers are included — training on this data… See the full description on the dataset page: https://huggingface.co/datasets/PrasannSinghal/ctc-suite-eval.ctc-cell-cycle-hela
CTC Cell Cycle Dataset
Cell Tracking Challenge (CTC) live-cell microscopy with derived cell cycle state
labels for 3-class temporal classification.
What's actually hosted
The repo name says hela for historical reasons. Currently hosted: Fluo-N2DH-GOWT1
(GFP-tagged Oct4 in mouse embryonic stem cells), which is what the milestone baseline
trained on. HeLa data may be added later under a hela/ prefix.
Sequence
Frames
Used as
01/
92
training
02/
92
held-out… See the full description on the dataset page: https://huggingface.co/datasets/DnaRnaProteins/ctc-cell-cycle-hela.ctcr_unity_rgb_seg_xyz_relativeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 75,
"total_frames": 62176,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:75"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JackySunUofT/ctcr_unity_rgb_seg_xyz_relative.ct-criterion-review
CT / Criterion Timing Review Dataset
A unified dataset of 6,241 proactive-assistant timing samples over 6,215
full-length videos (~128 GB), plus the complete human-review web platform
used to audit them.
Each sample pairs a video with a user request (e.g. "Walk me through
assembling this side table, and check that I align the leg joints correctly")
and a list of help points — the moments where a proactive assistant should
speak up, what it should say, and precisely when. The… See the full description on the dataset page: https://huggingface.co/datasets/LCZZZZ/ct-criterion-review.resd_ctc16000
RESD (CTC, 16 kHz)
RESD resampled to 16 kHz with wav2vec2 features precomputed.
How it was recorded
RESD was recorded in a studio by 20 voice actors. There was no script: the actors were not handed lines to read. Instead each actor in a pair was privately given an emotion to play, and the dialogue was improvised from there. So the words are spontaneous while the emotion is deliberate — which is the point, and also the limit. The label describes what the actor was… See the full description on the dataset page: https://huggingface.co/datasets/Aniemore/resd_ctc16000.HEVC_SDR_CTC
