CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Cybrpgs /stillalive-overlap-random-1000gated Stillalive random overlap sample This dataset contains 1,000 randomly selected audio chunks from the locally validated Stillalive dataset for which has_any_overlap is true and overlap_percentage is greater than zero. The complete source schema and embedded 48 kHz mono, 320 kbps MP3 audio are preserved. The sample contains 571 rows from finalized Part A and 429 rows from the current Part B merge. Sampling used deterministic reservoir sampling with seeds 20260920 (Part A) and… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-overlap-random-1000.audioautomatic-speech-recognition1K<n<10K0 likes14 downloads3d agoHugging Face02Cybrpgs /corpus5-proposed-no-overlap-random-1000gated Corpus5 proposed-rule unflagged random sample This manually gated dataset contains 1,000 uniformly randomly selected rows from the fixed 2,715,793-row prepared Corpus5 snapshot on mac02. A row is eligible only when both proposed checks are false: the full chunk interval does not intersect positive-duration diarization turns from two distinct speakers; and no speaker-change/no-change disagreement is detected at aligned adjacent words in transcript1 and transcript2 and propagated… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/corpus5-proposed-no-overlap-random-1000.audioautomatic-speech-recognition1K<n<10K0 likes14 downloads2d agoHugging Face03Cybrpgs /stillalive-proposed-no-overlap-random-1000gated Random sample of chunks not flagged by the proposed overlap rules This dataset contains 1,000 uniformly randomly selected chunks from the 1,408,153 chunks not flagged by either proposed overlap rule in a fixed 3,882,782-row completed Stillalive dataset snapshot, captured on 20 September 2026 at approximately 21:23 WAT. Selection used a single global reservoir across both dataset parts with seed 202609201. There were no additional filters for duration, text, language, quality… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-proposed-no-overlap-random-1000.audioautomatic-speech-recognition1K<n<10K0 likes11 downloads3d agoHugging Face04Cybrpgs /stillalive-proposed-no-overlap-random-1000-v2gated Second random sample of chunks not flagged by the proposed overlap rules This dataset contains a second, disjoint set of 1,000 uniformly randomly selected chunks from a fixed 3,882,782-row completed Stillalive dataset snapshot captured on 20 September 2026. The candidate population contains chunks for which both proposed flags are false: no two distinct ordinary-diarization speakers have positive-duration intersections with the full chunk interval; and no… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-proposed-no-overlap-random-1000-v2.audioautomatic-speech-recognition1K<n<10K0 likes10 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.