CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PrasannSinghal /ctc-suite-eval CTC suite eval ladders The 22-task corpus-tracking-capacity suite: per-task context ladders from 2k to 1M tokens, consumed by the ctc_suite task family on the prasann/ctc-suite branch of allenai/olmo-eval (ctc_nq:r64k, suites ctc:figure / ctc:xlong / ctc:r128k / ...). One config per task, one split per rung; each row is one unified-format example (documents + queries + answers + gold). Public release note (2026-08-14). Gold answers are included — training on this data… See the full description on the dataset page: https://huggingface.co/datasets/PrasannSinghal/ctc-suite-eval.tabular10K<n<100K0 likes3k downloads1mo agoHugging Face02JackySunUofT /ctcr_unity_rgb_seg_xyz_relativeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "unity", "total_episodes": 75, "total_frames": 62176, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:75" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JackySunUofT/ctcr_unity_rgb_seg_xyz_relative.imagerobotics10K<n<100K0 likes323 downloads5mo agoHugging Face03Aniemore /resd_ctc16000 RESD (CTC, 16 kHz) RESD resampled to 16 kHz with wav2vec2 features precomputed. How it was recorded RESD was recorded in a studio by 20 voice actors. There was no script: the actors were not handed lines to read. Instead each actor in a pair was privately given an emotion to play, and the dialogue was improvised from there. So the words are spontaneous while the emotion is deliberate — which is the point, and also the limit. The label describes what the actor was… See the full description on the dataset page: https://huggingface.co/datasets/Aniemore/resd_ctc16000.textaudio-classification1K<n<10K1 likes145 downloads2mo agoHugging Face04JackySunUofT /ctcr_unity_liquid_rgb_segThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "unity", "total_episodes": 75, "total_frames": 62176, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:75" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JackySunUofT/ctcr_unity_liquid_rgb_seg.imagerobotics10K<n<100K0 likes129 downloads5mo agoHugging Face05touringwithayo /nollywood-ctc-scored-ep3-hauwaaudion<1K1 likes90 downloads10d agoHugging Face06JackySunUofT /ctcr_unity_c1_rgb_seg_depth_100This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "unity", "total_episodes": 100, "total_frames": 22929, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:100" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JackySunUofT/ctcr_unity_c1_rgb_seg_depth_100.imagerobotics10K<n<100K0 likes81 downloads4mo agoHugging Face07bobboyms /phoneme-ctc-spanish-52h-noisyaudio10K<n<100K0 likes55 downloads9mo agoHugging Face08bobboyms /phoneme-ctc-english-60haudio10K<n<100K0 likes47 downloads9mo agoHugging Face09bobboyms /phoneme-ctc-english-41haudio10K<n<100K0 likes46 downloads11mo agoHugging Face10prakhya15 /hindi-conformer-ctc-segmentsaudio10K<n<100K0 likes44 downloads3d agoHugging Face11bobboyms /phoneme-ctc-english-60h-balanced Phoneme CTC — English 60h (Balanced & Normalized) A cleaned, normalized and phoneme-balanced version of bobboyms/phoneme-ctc-english-60h-noisy, for training phoneme recognition models (CTC) — e.g. as the native acoustic model behind pronunciation-feedback systems. What's different from the source dataset Label noise removed Roman numerals dropped — eSpeak reads ii/iv/… as "Roman two/four", producing labels that don't match the audio. Non-English phonemes dropped… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/phoneme-ctc-english-60h-balanced.audioautomatic-speech-recognition10K<n<100K0 likes40 downloads3mo agoHugging Face12akmalsultanov /usc_cleaned_ctc_filteredgatedaudio10K<n<100K0 likes39 downloads1y agoHugging Face13mi-rei /CT_complete Contains: TRIAL NAME BRIEF DRUG USED DRUG CLASS INDICATION TARGET THERAPY LEAD SPONSOR CRITERIA PRIMARY OUTCOME SECONDARY OUTCOME 1 SECONDARY OUTCOME 2 DOSAGE DESCRIPTION CONTROL DOSAGE DESCRIPTION TRIAL DESCRIPTION text10K<n<100K0 likes32 downloads3y agoHugging Face14nojima-kanta-ctc /yam_cup_bidirectional_v1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_yam_follower_robot", "total_episodes": 4, "total_frames": 2195, "total_tasks": 2, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:4" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/nojima-kanta-ctc/yam_cup_bidirectional_v1.tabularrobotics1K<n<10K0 likes16 downloads7mo agoHugging Face15aklemen /whisper-ctc-h2ttext100K<n<1M0 likes15 downloads1y agoHugging Face16bobboyms /phoneme-ctc-english-60h-noisyaudio10K<n<100K0 likes14 downloads9mo agoHugging Face17AhunInteligence /hubert_ctc_ftaudio1K<n<10K0 likes9 downloads8mo agoHugging Face18jmcochrane /SDS_FINAL_Noise_Level0.0001_NO_CTCtabularn<1K0 likes9 downloads5mo agoHugging Face19NancyT /wur_ctc_kln_scoredaudion<1K0 likes9 downloads3mo agoHugging Face20jmcochrane /SDS_FINAL_Noise_Level1e-05_NO_CTCtabularn<1K0 likes8 downloads5mo agoHugging Face21jmcochrane /SDS_FINAL_Noise_Level0.005_NO_CTCtabularn<1K0 likes8 downloads5mo agoHugging Face22d1shs0ap /sudoku-ctc-reasoning-processed-short-solutiontextn<1K0 likes7 downloads9mo agoHugging Face23jmcochrane /SDS_FINAL_Noise_Level0.0005_NO_CTCtabularn<1K0 likes7 downloads5mo agoHugging Face24IAmNotAnanth /sinhala-ctc-111hgated Sinhala ASR – Consolidated OpenSLR (SLR52) Dataset Summary This dataset is a consolidated and cleaned version of the Sinhala Automatic Speech Recognition (ASR) dataset from OpenSLR (SLR52). The original OpenSLR release distributes the data across multiple subsets (0–9, a–f). This repository merges all subsets into a single unified dataset containing approximately 111 hours of speech audio. Dataset Description Consolidation All OpenSLR SLR52… See the full description on the dataset page: https://huggingface.co/datasets/IAmNotAnanth/sinhala-ctc-111h.audioautomatic-speech-recognition100K<n<1M0 likes6 downloads9mo agoHugging Face25Purvaxxx /Marathi_CTC1K<n<10K0 likes5 downloads1y agoHugging Face26Purvaxxx /hindi_CTCtext1K<n<10K0 likes5 downloads1y agoHugging Face27jmcochrane /SDS_Sim_Noise_Level5e-05_NO_CTCtabularn<1K0 likes5 downloads6mo agoHugging Face28jmcochrane /SDS_Sim_Noise_Level1e-05_NO_CTCtabularn<1K0 likes4 downloads6mo agoHugging Face29jmcochrane /SDS_FINAL_Noise_Level5e-05_NO_CTCtabularn<1K0 likes4 downloads5mo agoHugging Face30jmcochrane /SDS_FINAL_Noise_Level0.001_NO_CTCtabularn<1K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.