datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AyaVisionBench
Dataset Card for Aya Vision Benchmark
Dataset Details
The Aya Vision Benchmark is designed to evaluate vision-language models in real-world multilingual scenarios. It spans 23 languages and 9 distinct task categories, with 15 samples per category, resulting in 135 image-question pairs per language.
Each question requires visual context for the answer and covers languages that half of the world's population speaks, making this dataset particularly suited for… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/AyaVisionBench.afri-aya-vision
Afri-Aya Vision: Restructured for Multimodal & Adaption Fine-Tuning
This dataset is a restructured, multimodal Vision-Language (VLM) adaptation of CohereLabsCommunity/afri-aya (Giving Sight to African LLMs).
Why This Restructured Version?
The original Afri-Aya dataset stores multiple question-and-answer pairs per image inside a nested list column (qa_pairs). Fine-tuning platforms (such as Adaption, Unsloth, LLaVA, and standard VLM training harnesses) require:
1… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/afri-aya-vision.SEA-AyaVisionBenchVLMEvalKit_AyaVisionBench
Aya Vision Bench for VLMEvalKit
Original dataset: ported to VLMEvalKit
Multilingual dataset spans 23 languages and 9 distinct task categories, with 15 samples per category, resulting in 135 image-question pairs per language.
Original dataset row:
{'image': [PIL.Image],
'image_source': 'VisText',
'image_source_category': 'Chart/figure understanding',
'index' : '17'
'question': 'If the top three parties by vote percentage formed a coalition, what percentage of the total votes… See the full description on the dataset page: https://huggingface.co/datasets/timothycdc/VLMEvalKit_AyaVisionBench.adaption-afri-aya-vision-restructured
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-afri_aya_vision_restructured
This dataset is a restructured multimodal Vision-Language collection derived from Afri-Aya, specifically formatted for fine-tuning models on African languages. It contains image-based question-answer pairs covering diverse topics such as traditional attire, instruments, and cultural practices across languages like Kinyarwanda, Yoruba, Zulu… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/adaption-afri-aya-vision-restructured.ro_ayavisionbench
Dataset Description
AyaVisionBench is designed to evaluate vision-language models in real-world multilingual scenarios.
Here we provide the Romanian version of AyaVisionBench augmented with human curated answers. This dataset is used as a benchmark and is part of the evaluation protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@article{dash2025aya,
title={Aya vision:… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_ayavisionbench.
