datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
visual-qa-llama-format
Open Paws Visual Qa Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Multimodal Data
Format: JSONL (JSON Lines)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/visual-qa-llama-format.visual_qa_histograms
What Lies Beneath: A Call for Distribution-based Visual Question & Answer Datasets
Publication: JCDL 2025 Website and on arXiv
GitHub Repo: ReadingTimeMachine/LLM_VQA_JCDL2025
This is a histogram-based dataset for visual question and answer (VQA) with humans and large language/multimodal models (LMMs).
Data contains synthetically generated single-panel histograms images, data used to create histograms, bounding box data for titles, axis and tick labels, and… See the full description on the dataset page: https://huggingface.co/datasets/ReadingTimeMachine/visual_qa_histograms.UMIE-Visual-QAmath-visual-qa-cunaruto-visual-qareddit-roastme-visual-qa
