datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset
Rapidata Image Generation Coherence Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.FluxHands-FingerCount
Random Sample
Citation
@misc{FluxHandsFingerCount,
title = {FluxHands-FingerCount Dataset},
author = {Taesiri, Mohammad Reza and Ghotbizadeh, Marjan and Mirabolghasemi, Pejman},
year = {2025},
howpublished = {Hugging Face Datasets},
url = {https://huggingface.co/datasets/taesiri/FluxHands-FingerCount},
}
V-FLUTE
Description
Large Vision-Language models (VLMs) have demonstrated strong reasoning capabilities in tasks requiring a fine-grained understanding of literal images and text, such as visual question-answering or visual entailment. However, there has been little exploration of these models' capabilities when presented with images and captions containing figurative phenomena such as metaphors or humor, the meaning of which is often implicit. To close this gap, we propose a new task and… See the full description on the dataset page: https://huggingface.co/datasets/ColumbiaNLP/V-FLUTE.
