datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TIC-TALK
TIC-TALK — Timing In stand-up Comedy: Text, Audio, Laughter, Kinesics
Time-aligned text, audio, and vision features from 90 stand-up specials
(2015–2024). Show identities are anonymized (SHOW_XXXX); no source audio,
video, or transcripts are distributed.
Dataset contains:
Shows
90
Total runtime
≈ 94 h (mean 63 min / show)
Rows (60 s blocks)
5 416
Topics (BERTopic)
24
Video frames (1 fps)
322 973 (22 % full-body)
Laughter coverage (mean)
17.8 % of block… See the full description on the dataset page: https://huggingface.co/datasets/ENC-PSL/TIC-TALK.PSL
Protein Subcellular Localization Extraction Task
Starling Query: Find all proteins mentioned in PubMed for which subcellular localization has been reported.
Schema Specification
localization_sites
Type: string
Description: Subcellular compartment(s) where the protein is stated to localize/accumulate/be retained/targeted/enriched; semicolon-separate multiple sites within the same context (e.g., "Golgi apparatus; endoplasmic reticulum").… See the full description on the dataset page: https://huggingface.co/datasets/starling-labs/PSL.evahan-ultraglyph
Chinese OCR Line Dataset
This repository contains the synthetic dataset used by the ENCHANTeam team for the EvaHan 2026 competition on OCR/HTR of Chinese documents.
Final ranking of the team: 3rd.
UltraGlyph data consist in a automatic mix of real data provided by the organizers, for the closed modality of the competition (no external dataset authorized).
Resource
Link
Preprint
HAL
Code
GitHub
Competition
EvaHan @ LREC 2026
Dataset description
Pairs… See the full description on the dataset page: https://huggingface.co/datasets/ENC-PSL/evahan-ultraglyph.sign-psl-13bBSICLE-medieval-illumination-folio-bin-class-dataset
BSICLE Medieval Folio Illumination Dataset
This dataset contains 1,484 medieval and early modern folio images, dating approximately from the 7th to the mid-17th century. annotated for binary image classification: whether a folio contains illumination or not.
The dataset is intended to train and evaluate
lightweight computer vision models for the fast
detection of illuminated folios in medieval manuscript
corpora, especially in IIIF-based heritage and
research workflows.
Check… See the full description on the dataset page: https://huggingface.co/datasets/ENC-PSL/BSICLE-medieval-illumination-folio-bin-class-dataset.psle-math
