datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
celebahq_512_id_clusters
celebahq_512 with SRK identity labels
Summary
This dataset is a derived version of jxie/celeba-hq. It keeps the original image set and adds automatically generated identity-group labels derived from face-embedding clustering.
As explained in our experimental setup, we use CelebA-HQ from Karras et al. (2018), specifically the Hugging Face snapshot at revision 7ecc6a45edfb5483ccf2f7df1035d298ffe7c76b. The referenced CelebA-HQ version provides gender labels but no identity… See the full description on the dataset page: https://huggingface.co/datasets/edgarcancinoe/celebahq_512_id_clusters.doc-content-clustering-740
Document Content-Clustering Benchmark (12 classes, 740 items)
A small, curated benchmark for clustering documents by their content topic
(not by their visual form/layout). Each item is a single document page provided
as an image plus two text views (a VLM description and OCR markdown), with a
ground-truth content class.
The set is intentionally "tangle-stripped": 27 borderline items whose content
sits ambiguously between two classes were removed from a larger 907-item pool to… See the full description on the dataset page: https://huggingface.co/datasets/langminer/doc-content-clustering-740.synchro-April2025-cluster-labeled-highMag
IFCB Plankton Labeled (Cluster-Sorted)
This dataset contains labeled images of phytoplankton collected with the Planktivore Imaging System. Images were preprocessed with a zero-padding and resized to the standard size used for ViT_b_16
The dataset was originally constructed by clustering unlabeled ROI images using deep features from a ViT model.Clusters were then saved locally and manually curated into taxonomic labels and higher-order groups.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/patcdaniel/synchro-April2025-cluster-labeled-highMag.
