datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ctc-suite-eval
CTC suite eval ladders
The 22-task corpus-tracking-capacity suite: per-task context ladders from 2k to 1M tokens,
consumed by the ctc_suite task family on the prasann/ctc-suite branch of allenai/olmo-eval
(ctc_nq:r64k, suites ctc:figure / ctc:xlong / ctc:r128k / ...). One config per task, one
split per rung; each row is one unified-format example (documents + queries + answers + gold).
Public release note (2026-08-14). Gold answers are included — training on this data… See the full description on the dataset page: https://huggingface.co/datasets/PrasannSinghal/ctc-suite-eval.ctc-cell-cycle-hela
CTC Cell Cycle Dataset
Cell Tracking Challenge (CTC) live-cell microscopy with derived cell cycle state
labels for 3-class temporal classification.
What's actually hosted
The repo name says hela for historical reasons. Currently hosted: Fluo-N2DH-GOWT1
(GFP-tagged Oct4 in mouse embryonic stem cells), which is what the milestone baseline
trained on. HeLa data may be added later under a hela/ prefix.
Sequence
Frames
Used as
01/
92
training
02/
92
held-out… See the full description on the dataset page: https://huggingface.co/datasets/DnaRnaProteins/ctc-cell-cycle-hela.ctcr_unity_rgb_seg_xyz_relativeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 75,
"total_frames": 62176,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:75"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JackySunUofT/ctcr_unity_rgb_seg_xyz_relative.ct-criterion-review
CT / Criterion Timing Review Dataset
A unified dataset of 6,241 proactive-assistant timing samples over 6,215
full-length videos (~128 GB), plus the complete human-review web platform
used to audit them.
Each sample pairs a video with a user request (e.g. "Walk me through
assembling this side table, and check that I align the leg joints correctly")
and a list of help points — the moments where a proactive assistant should
speak up, what it should say, and precisely when. The… See the full description on the dataset page: https://huggingface.co/datasets/LCZZZZ/ct-criterion-review.resd_ctc16000
RESD (CTC, 16 kHz)
RESD resampled to 16 kHz with wav2vec2 features precomputed.
How it was recorded
RESD was recorded in a studio by 20 voice actors. There was no script: the actors were not handed lines to read. Instead each actor in a pair was privately given an emotion to play, and the dialogue was improvised from there. So the words are spontaneous while the emotion is deliberate — which is the point, and also the limit. The label describes what the actor was… See the full description on the dataset page: https://huggingface.co/datasets/Aniemore/resd_ctc16000.HEVC_SDR_CTCChangeMore-prompt-injection-eval
ChangeMore-prompt-injection-eval
This dataset is designed to support the evaluation of prompt injection detection capabilities in large language models (LLMs).
To address the lack of Chinese-language prompt injection attack samples, we have developed a systematic data generation algorithm that automatically produces a large volume of high-quality attack samples. These samples significantly enrich the security evaluation ecosystem for LLMs, especially in the Chinese context.
The… See the full description on the dataset page: https://huggingface.co/datasets/CTCT-CT2/ChangeMore-prompt-injection-eval.ctcr_unity_liquid_rgb_segThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 75,
"total_frames": 62176,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:75"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JackySunUofT/ctcr_unity_liquid_rgb_seg.nollywood-ctc-scored-ep3-hauwactcr_unity_c1_rgb_seg_depth_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 100,
"total_frames": 22929,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JackySunUofT/ctcr_unity_c1_rgb_seg_depth_100.phoneme-ctc-spanish-52h-noisynlp-sentiment-multimodal3
NLP Sentiment Multimodal3 Data Notes
Dataset summary
A documented NLP Sentiment data-preparation workflow for Multimodal3 records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/ctchen5674/nlp-sentiment-multimodal3.phoneme-ctc-english-60hindia-ctc-to-in-hand-salary-2026
India CTC-to-In-Hand Salary Dataset 2026
This dataset models how annual Cost to Company (CTC) converts into annual net cash and regular monthly net pay for salaried employees in India.
It covers CTC bands from ₹5 lakh to ₹50 lakh under:
0%, 50% and 100% target-variable payout scenarios
capped and full-basic-wage provident-fund models
the Income-tax Act, 2025 provisions applicable to tax year 2026–27
Why this dataset exists
CTC, recurring monthly bank credit and… See the full description on the dataset page: https://huggingface.co/datasets/PaisaSamajhResearch/india-ctc-to-in-hand-salary-2026.phoneme-ctc-english-41hPrompt-CHIP-CTCusc_cleaned_ctc_filteredtranslategemma-4b-it-Q4_K_M-GGUF
TranslateGemma 4B IT Q4_K_M GGUF
This repository contains a GGUF Q4_K_M conversion of Google TranslateGemma 4B IT.
Model information
Base model: google/translategemma-4b-it
Source revision: 10042cb0e6e7fdce748996a71dc3dc432a4e0c89
llama.cpp revision used in the project validation: 2048b5913d51beab82dfe29955f9008130b936c0
Artifact filename: translategemma-4b-it-Q4_K_M.gguf
Size: 2489909120 bytes
SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/ctc88haha/translategemma-4b-it-Q4_K_M-GGUF.uzvoice_cleaned_ctc_filteredCT_complete
Contains:
TRIAL NAME
BRIEF
DRUG USED
DRUG CLASS
INDICATION
TARGET
THERAPY
LEAD SPONSOR
CRITERIA
PRIMARY OUTCOME
SECONDARY OUTCOME 1
SECONDARY OUTCOME 2
DOSAGE DESCRIPTION
CONTROL DOSAGE DESCRIPTION
TRIAL DESCRIPTION
CT-Chat-Docsct_cases_samplesphoneme-ctc-english-60h-balanced
Phoneme CTC — English 60h (Balanced & Normalized)
A cleaned, normalized and phoneme-balanced version of
bobboyms/phoneme-ctc-english-60h-noisy,
for training phoneme recognition models (CTC) — e.g. as the native acoustic
model behind pronunciation-feedback systems.
What's different from the source dataset
Label noise removed
Roman numerals dropped — eSpeak reads ii/iv/… as "Roman two/four",
producing labels that don't match the audio.
Non-English phonemes dropped… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/phoneme-ctc-english-60h-balanced.icefall-vsr-grid-conformer-ctc2-bpe-58-2026-04-29ctc-seed-pools
CTC seed pools
Serialized corpus pools for the CTC long-context suite's data generators
(allenai/OLMo-core, branch prasann/ctc,
pip package ctc/). Each file is the output of the one build step that needs heavy machinery — a
GPU cross-encoder, a pyserini/Lucene index, an LLM mining run, or a multi-gigabyte download —
captured once, so that anyone can build train and eval data at any context scale (2k to 10M+
tokens per example) on a bare pip install: no GPU, no Java, no API key.… See the full description on the dataset page: https://huggingface.co/datasets/PrasannSinghal/ctc-seed-pools.m2m100-418m-fp16-merged-onnx-ios-package
M2M100 418M FP16 Merged ONNX iOS Package
This repository contains an ONNX FP16 merged runtime package converted from
facebook/m2m100_418M for use
in an offline iOS translation app.
This is a converted runtime package. It is not the original unmodified PyTorch
model checkpoint published by Meta/Facebook.
This repository is not endorsed by Meta/Facebook.
Package contents
The archive m2m100-418m-fp16-merged-ios.zip contains one top-level folder… See the full description on the dataset page: https://huggingface.co/datasets/ctc88haha/m2m100-418m-fp16-merged-onnx-ios-package.yam_cup_bidirectional_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_yam_follower_robot",
"total_episodes": 4,
"total_frames": 2195,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/nojima-kanta-ctc/yam_cup_bidirectional_v1.whisper-ctc-h2tctchat_evaluationphoneme-ctc-english-60h-noisy
