datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human-coherence-preferences-images
Rapidata Image Generation Coherence Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated coherence datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-coherence-preferences-images.human-alignment-preferences-images
Rapidata Image Generation Alignment Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated alignment datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-alignment-preferences-images.thai_famous_people_images_dataset
Thai Famous People Image Dataset
Dataset Description
The Thai Famous People Image Dataset is a collection of images and descriptions of famous Thai personalities. This dataset is designed to provide a comprehensive resource for researchers, developers, and enthusiasts interested in Thai culture, history, and notable figures. The data was extracted from the Thai Wikipedia dump in September 2024, ensuring up-to-date and relevant information.
Maintainer
Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/iapp/thai_famous_people_images_dataset.Hard_images_for_VLMsocr-vqa-200k_imagesImage collections for OCR-VQA-200K.
Image size: 208,467.
CuPer_Images
CuPer_Images
Resumen del Dataset
CuPer_Images es un dataset multimodal especializado en culturas precolombinas de Perú, desarrollado por NovaIA, el laboratorio de inteligencia artificial de Grupo Neura. Este dataset fue utilizado como parte del entrenamiento de Amaru, un modelo de lenguaje de propósito general con conocimientos profundos en las civilizaciones ancestrales del Perú.
Información del Dataset
Tamaño total: 4,247 muestras
División: 3,396 muestras… See the full description on the dataset page: https://huggingface.co/datasets/NovaIALATAM/CuPer_Images.leuven-images
