datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AMALIA-VL-SFT-Dataset
AMALIA-VL-Training-Dataset
Dataset Description
This dataset is provided as part of the AMALIA project.
This is the vision+language training mix for AMALIA-VL-SFT. Each subset is one
source dataset in the mix, each with a single train split. The only datasets that are absent from this mix are those that derive directly from the core LLM training mix, and can be found in the AMALIA-LLM Post Training Collection.
Example usage:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-VL-SFT-Dataset.SEED-Bench-PT
SEED-Bench-PT
European Portuguese (pt-PT) machine translation of SEED-Bench, a multiple-choice benchmark spanning multiple dimensions of multimodal comprehension.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/SEED-Bench
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/SEED-Bench-PT.InfographicVQA-PT
InfographicVQA-PT
European Portuguese (pt-PT) machine translation of InfographicVQA, a visual question answering dataset over infographics that combine text, graphics, and data visualizations.
Translated from the original English validation split (InfographicVQA subset) using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/DocVQA (InfographicVQA subset)
Note: This dataset is machine translated and may contain translation errors or artifacts.… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/InfographicVQA-PT.COCO-Caption2017-PT
COCO-Caption2017-PT
European Portuguese (pt-PT) machine translation of COCO Captions 2017, an image captioning dataset of everyday scenes with human-written captions.
Translated from the original English val split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/COCO-Caption2017
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/COCO-Caption2017-PT.CARAVELA
CARAVELA
CARAVELA is a multimodal benchmark for evaluating the Portuguese cultural knowledge of large vision-language models (LVLMs). The official benchmark language is exclusively European Portuguese (pt-PT).
Motivation
Modern LVLMs excel at general-purpose vision-language tasks, but their performance drops sharply on localized, culturally specific content that is underrepresented in global training data. Portugal has a rich cultural heritage —… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/CARAVELA.TextVQA-PT
TextVQA-PT
European Portuguese (pt-PT) machine translation of TextVQA, a visual question answering dataset that requires reading and reasoning about text in images.
Translated from the original English validation split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/textvqa
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/TextVQA-PT.MMMU-Pro-PT
MMMU-Pro-PT
European Portuguese (pt-PT) machine translation of MMMU-Pro, a more robust and challenging version of MMMU for college-level, multi-discipline multimodal reasoning.
Translated from the original English test split (standard (10 options) subset) using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/MMMU/MMMU_Pro (standard (10 options) subset)
Note: This dataset is machine translated and may contain translation errors or artifacts.
This… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MMMU-Pro-PT.RealWorldQA-PT
RealWorldQA-PT
European Portuguese (pt-PT) machine translation of RealWorldQA, a benchmark of real-world spatial understanding questions from images captured in physical environments.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/RealWorldQA
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/RealWorldQA-PT.MATH-Vision-PT
MATH-Vision-PT
European Portuguese (pt-PT) machine translation of MATH-Vision, a benchmark of competition-level mathematics problems presented in visual contexts.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/MathLLMs/MathVision
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MATH-Vision-PT.EmbSpatial-Bench-PT
EmbSpatial-Bench-PT
European Portuguese (pt-PT) machine translation of EmbSpatial-Bench, a benchmark for evaluating spatial reasoning from an embodied, egocentric perspective.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/FlagEval/EmbSpatial-Bench
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/EmbSpatial-Bench-PT.MMMU-PT
MMMU-PT
European Portuguese (pt-PT) machine translation of MMMU, a benchmark of college-level, multi-discipline questions requiring expert-level multimodal understanding and reasoning.
Translated from the original English validation split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/MMMU
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MMMU-PT.OCRBench-PT
OCRBench-PT
European Portuguese (pt-PT) machine translation of OCRBench, a benchmark evaluating the optical character recognition and text understanding abilities of multimodal models.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/echo840/OCRBench
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/OCRBench-PT.MMStar-PT
MMStar-PT
European Portuguese (pt-PT) machine translation of MMStar, a curated multimodal benchmark of vision-indispensable, balanced challenge samples.
Translated from the original English val split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/Lin-Chen/MMStar
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in amalia-vl-eval, a… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MMStar-PT.ChartQA-PT
ChartQA-PT
European Portuguese (pt-PT) machine translation of ChartQA, a visual question answering dataset that requires reasoning over charts and plots.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/ChartQA
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in amalia-vl-eval, a… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/ChartQA-PT.DocVQA-PT
DocVQA-PT
European Portuguese (pt-PT) machine translation of DocVQA, a visual question answering dataset over scanned document images.
Translated from the original English validation split (DocVQA subset) using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/DocVQA (DocVQA subset)
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/DocVQA-PT.POPE-PT
POPE-PT
European Portuguese (pt-PT) machine translation of POPE, a benchmark for evaluating object hallucination in vision-language models through yes/no polling questions.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/POPE
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/POPE-PT.AI2D-PT
AI2D-PT
European Portuguese (pt-PT) machine translation of AI2D, a multiple-choice visual question answering dataset over grade-school science diagrams.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/ai2d
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in amalia-vl-eval, a… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AI2D-PT.AMALIA-VL-DPO-Dataset
AMALIA-VL-DPO Dataset
Dataset Description
This dataset is provided as part of the AMALIA project.
Direct Preference Optimization (DPO) data used for AMALIA-VL. Each subset is one source in
the mix, with vl_preference_200k being fully derived from the SFT mix. Every row is a preference triplet:
column
description
prompt
normalized [{role, content}] turns; <image> marks image position (multimodal subsets)
chosen
preferred assistant response… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-VL-DPO-Dataset.MME-PT
MME-PT
European Portuguese (pt-PT) machine translation of MME, a multimodal benchmark measuring perception and cognition abilities through yes/no questions.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/MME
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in amalia-vl-eval, a… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MME-PT.
