datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ctc-suite-eval
CTC suite eval ladders
The 22-task corpus-tracking-capacity suite: per-task context ladders from 2k to 1M tokens,
consumed by the ctc_suite task family on the prasann/ctc-suite branch of allenai/olmo-eval
(ctc_nq:r64k, suites ctc:figure / ctc:xlong / ctc:r128k / ...). One config per task, one
split per rung; each row is one unified-format example (documents + queries + answers + gold).
Public release note (2026-08-14). Gold answers are included — training on this data… See the full description on the dataset page: https://huggingface.co/datasets/PrasannSinghal/ctc-suite-eval.ctcr_unity_rgb_seg_xyz_relativeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 75,
"total_frames": 62176,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:75"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JackySunUofT/ctcr_unity_rgb_seg_xyz_relative.resd_ctc16000
RESD (CTC, 16 kHz)
RESD resampled to 16 kHz with wav2vec2 features precomputed.
How it was recorded
RESD was recorded in a studio by 20 voice actors. There was no script: the actors were not handed lines to read. Instead each actor in a pair was privately given an emotion to play, and the dialogue was improvised from there. So the words are spontaneous while the emotion is deliberate — which is the point, and also the limit. The label describes what the actor was… See the full description on the dataset page: https://huggingface.co/datasets/Aniemore/resd_ctc16000.ctcr_unity_liquid_rgb_segThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 75,
"total_frames": 62176,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:75"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JackySunUofT/ctcr_unity_liquid_rgb_seg.nollywood-ctc-scored-ep3-hauwactcr_unity_c1_rgb_seg_depth_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 100,
"total_frames": 22929,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JackySunUofT/ctcr_unity_c1_rgb_seg_depth_100.phoneme-ctc-spanish-52h-noisyphoneme-ctc-english-60hphoneme-ctc-english-41hhindi-conformer-ctc-segmentsphoneme-ctc-english-60h-balanced
Phoneme CTC — English 60h (Balanced & Normalized)
A cleaned, normalized and phoneme-balanced version of
bobboyms/phoneme-ctc-english-60h-noisy,
for training phoneme recognition models (CTC) — e.g. as the native acoustic
model behind pronunciation-feedback systems.
What's different from the source dataset
Label noise removed
Roman numerals dropped — eSpeak reads ii/iv/… as "Roman two/four",
producing labels that don't match the audio.
Non-English phonemes dropped… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/phoneme-ctc-english-60h-balanced.usc_cleaned_ctc_filteredCT_complete
Contains:
TRIAL NAME
BRIEF
DRUG USED
DRUG CLASS
INDICATION
TARGET
THERAPY
LEAD SPONSOR
CRITERIA
PRIMARY OUTCOME
SECONDARY OUTCOME 1
SECONDARY OUTCOME 2
DOSAGE DESCRIPTION
CONTROL DOSAGE DESCRIPTION
TRIAL DESCRIPTION
yam_cup_bidirectional_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_yam_follower_robot",
"total_episodes": 4,
"total_frames": 2195,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/nojima-kanta-ctc/yam_cup_bidirectional_v1.whisper-ctc-h2tphoneme-ctc-english-60h-noisyhubert_ctc_ftSDS_FINAL_Noise_Level0.0001_NO_CTCwur_ctc_kln_scoredSDS_FINAL_Noise_Level1e-05_NO_CTCSDS_FINAL_Noise_Level0.005_NO_CTCsudoku-ctc-reasoning-processed-short-solutionSDS_FINAL_Noise_Level0.0005_NO_CTCsinhala-ctc-111h
Sinhala ASR – Consolidated OpenSLR (SLR52)
Dataset Summary
This dataset is a consolidated and cleaned version of the Sinhala Automatic Speech Recognition (ASR) dataset from OpenSLR (SLR52).
The original OpenSLR release distributes the data across multiple subsets (0–9, a–f).
This repository merges all subsets into a single unified dataset containing approximately 111 hours of speech audio.
Dataset Description
Consolidation
All OpenSLR SLR52… See the full description on the dataset page: https://huggingface.co/datasets/IAmNotAnanth/sinhala-ctc-111h.Marathi_CTChindi_CTCSDS_Sim_Noise_Level5e-05_NO_CTCSDS_Sim_Noise_Level1e-05_NO_CTCSDS_FINAL_Noise_Level5e-05_NO_CTCSDS_FINAL_Noise_Level0.001_NO_CTC
