datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AMALIA-VL-SFT-Dataset
AMALIA-VL-Training-Dataset
Dataset Description
This dataset is provided as part of the AMALIA project.
This is the vision+language training mix for AMALIA-VL-SFT. Each subset is one
source dataset in the mix, each with a single train split. The only datasets that are absent from this mix are those that derive directly from the core LLM training mix, and can be found in the AMALIA-LLM Post Training Collection.
Example usage:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-VL-SFT-Dataset.SEED-Bench-PT
SEED-Bench-PT
European Portuguese (pt-PT) machine translation of SEED-Bench, a multiple-choice benchmark spanning multiple dimensions of multimodal comprehension.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/SEED-Bench
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/SEED-Bench-PT.CorEGe-PT
CorEGe-PT: Corpus do Estudo Geral - Portuguese
CorEGe-PT is a large-scale corpus of academic texts written in Portuguese (mainly European Portuguese), extracted from Estudo Geral, the institutional repository of the University of Coimbra. It contains over 34,000 documents and approximately 1 billion tokens, making it the largest available corpus of its kind for the Portuguese language.
This dataset is designed to support linguistic research (Academic Discourse Studies) and the… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/CorEGe-PT.bigbenchhard-mt-pt
BBH-PT (Big-Bench Hard)
Portuguese machine translation of BIG-Bench Hard, a challenging subset of the BIG-Bench benchmark covering diverse reasoning tasks.
Translated using a Finetuned GemmaX2-9B for pt-PT with rule-based adaptations.
Note: Some tasks (e.g., hyperbaton) are not translated as they do not transfer meaningfully to Portuguese.
Original Dataset: https://github.com/suzgunmirac/BIG-Bench-Hard
Note: This dataset is machine translated and may contain… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/bigbenchhard-mt-pt.pt_exams
PHEB - Portuguese High School Exams MCQ
MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum.
For more details, see the PHEB paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.alba_mcq
ALBA MCQ
Multiple Choice Version of the ALBA benchmark, a Portuguese language benchmark for proficiency in pt-PT linguistic-related tasks.
For more details, see the ALBA paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/alba_mcq.hendrycks-math-ptpt
Hendrycks Math MT-PT
Portuguese translated mathematics problems covering algebra, geometry, number theory, and more.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/EleutherAI/hendrycks_math
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/hendrycks-math-ptpt.InfographicVQA-PT
InfographicVQA-PT
European Portuguese (pt-PT) machine translation of InfographicVQA, a visual question answering dataset over infographics that combine text, graphics, and data visualizations.
Translated from the original English validation split (InfographicVQA subset) using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/DocVQA (InfographicVQA subset)
Note: This dataset is machine translated and may contain translation errors or artifacts.… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/InfographicVQA-PT.AMALIA-LLM-0626-SFT-Dataset
AMALIA LLM Supervised Finetuning Dataset
Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training.
Base Data Mix
This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table:
Dataset
Count
amalia-llm/persona_math
63,731… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-0626-SFT-Dataset.persona_nemotron
Persona Nemotron PT Datasets
This is a collection of Portuguese synthetic datasets, consisting of 3 datasets, one with general questions from varied topics, one with math questions, and one with instruction-following requests.
The prompts were generated using an approach similar to PersonaHub, with a translated version of Nemotron Personas. Both prompts and answers were generated using Gemma 3-27B.
This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_nemotron.COCO-Caption2017-PT
COCO-Caption2017-PT
European Portuguese (pt-PT) machine translation of COCO Captions 2017, an image captioning dataset of everyday scenes with human-written captions.
Translated from the original English val split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/COCO-Caption2017
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/COCO-Caption2017-PT.CARAVELA
CARAVELA
CARAVELA is a multimodal benchmark for evaluating the Portuguese cultural knowledge of large vision-language models (LVLMs). The official benchmark language is exclusively European Portuguese (pt-PT).
Motivation
Modern LVLMs excel at general-purpose vision-language tasks, but their performance drops sharply on localized, culturally specific content that is underrepresented in global training data. Portugal has a rich cultural heritage —… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/CARAVELA.amalia-sft
AMALIA SFT
Standardized European Portuguese supervised fine-tuning data from AMALIA.
Each config has deterministic train and validation splits and a common messages schema.
Exact duplicate and invalid conversations are removed. Source provenance is retained in every row.
Counts
Config
Train
Validation
pt_persona_instruction
8902
182
pt_nemotron_instruction
4394
90
pt_nemotron_general
69304
1415
pt_wikipedia
95592
1951
pt_culture
88058
1798… See the full description on the dataset page: https://huggingface.co/datasets/duarteocarmo/amalia-sft.PT-Culture_Data
Portuguese Cultural SFT Dataset
A supervised fine-tuning dataset for European Portuguese (PT-PT) cultural knowledge, built to teach models the traditions, figures, places, and expressions of Portuguese culture.
The dataset contains 216,832 examples across two versions in conversational SFT format, organised into ten cultural domains:
Domain
v1 (95,815)
v2 (121,017)
Total
Personalities
36,370
62,788
99,158
Audiovisual
15,786
26,857
42,643
Heritage
8,020… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/PT-Culture_Data.smoltalk2_everyday_conv_pt
SMOL Everyday Conversation PT
This dataset consists of a Portuguese version of the smoltalk_smollm3_everyday_conversations_no_think split of HuggingFaceTB/smoltalk2. The first two user turns were translated as well as the first assistant turn, the continuation of the conversation was generated using Gemma 3-27B.
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smoltalk2_everyday_conv_pt.wikipedia-rag
Wikipedia RAG
This dataset consists of raw wikipedia pages with one question and answer for each page. The answers were generated using Gemma 3-27B.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title = "{AMALIA}: A Fully Open Large Language Model for… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wikipedia-rag.saudade-pt
SAUDADE
Portuguese benchmark for temporal reasoning and understanding of temporal relationships across Portuguese events.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title = "{AMALIA}: A Fully Open… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/saudade-pt.amalia-Dolci-Instruct-SFT
AMALIA Dolci-Instruct-SFT
Version of the allenai/Dolci-Instruct-SFT dataset used in the ramp down phase of AMALIA's Supervised Fine-Tuning stage.
This dataset was developed by sampling the highest quality entries of the selected splits, and translating part of those entries to European Portuguese. Both the quality classification and translation were done using google/gemma-4-31B-it.
This dataset went through a processing pipeline to:
Remove entries that reference… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Dolci-Instruct-SFT.persona_instruction_following
Persona Instruction Following Datasets
This is a synthetic instruction-following dataset, available in two configs: full and filtered. Each config contains two language splits, English (en) and Portuguese (pt). The filtered version keeps only the higher-quality examples (quality score 5).
The prompts were generated using an approach similar to PersonaHub, with a translated version of proj-persona/PersonaHub. Both prompts and answers were generated using Gemma 3-27B.… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_instruction_following.wildguardmix-ptpt
WildGuardTest-PT
Portuguese machine translation of WildGuardTest, a benchmark for evaluating safety guardrails in language models.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/walledai/WildGuardTest
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wildguardmix-ptpt.amalia-Nemotron-SFT-Math-v3
AMALIA Nemotron-SFT-Math-v3
Version of the nvidia/Nemotron-SFT-Math-v3 dataset used in the AMALIA's Supervised Fine-Tuning stage.
This dataset went through a processing pipeline to:
Remove entries with tool usage;
Remove entries that reference other LLMs or research labs;
Removed the reasoning_content field;
Drop extra unnecessary columns;
Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Math-v3
This dataset is provided as part of the AMALIA… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SFT-Math-v3.alba
ALBA
ALBA is a Portuguese language benchmark for proficiency in pt-PT linguistic-related tasks.
For more details, see the ALBA paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title = "{AMALIA}: A… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/alba.DPO-Dataset
AMALIA DPO Dataset
This is the DPO (preference optimization) dataset used to train AMALIA-DPO.
It is a mix of preference pairs, mainly in European Portuguese and English, covering general conversation, instruction following, math, and safety. These pairs come from different sources, including prompts from the SFT mix, responses generated by different models, including an early version of the model, and some public datasets.
This dataset is provided as part of the AMALIA project… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/DPO-Dataset.amalia-Nemotron-Instruction-Following-Chat-v1
AMALIA Nemotron-Instruction-Following-Chat-v1
Version of the nvidia/Nemotron-Instruction-Following-Chat-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage.
This dataset went through a processing pipeline to:
Remove entries where the capability_target field was 'chat' from the chat_if split;
Remove entries that reference other LLMs or research labs;
Remove entries that contained the string '/imagine prompt:';
Removed the reasoning_content field;
Original… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-Instruction-Following-Chat-v1.TextVQA-PT
TextVQA-PT
European Portuguese (pt-PT) machine translation of TextVQA, a visual question answering dataset that requires reading and reasoning about text in images.
Translated from the original English validation split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/textvqa
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/TextVQA-PT.MMMU-Pro-PT
MMMU-Pro-PT
European Portuguese (pt-PT) machine translation of MMMU-Pro, a more robust and challenging version of MMMU for college-level, multi-discipline multimodal reasoning.
Translated from the original English test split (standard (10 options) subset) using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/MMMU/MMMU_Pro (standard (10 options) subset)
Note: This dataset is machine translated and may contain translation errors or artifacts.
This… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MMMU-Pro-PT.hermes3_special_system_prompts
Hermes 3 SFT Special System Prompts
Version of the NousResearch/Hermes-3-Dataset dataset used in AMALIA's Supervised Fine-Tuning stage.
This subset was manually selected in order to retain only entries containing custom system prompts that substantially modify the model’s behavior. It were also built splits with the highest quality entries and a translated part of those entries to European Portuguese. Both the quality classification and translation were done using… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/hermes3_special_system_prompts.piqa-mt-pt
PIQA-PT
Portuguese machine translation of PIQA (Physical Interaction QA), a benchmark for physical commonsense reasoning.
Translated using a Finetuned GemmaX2-9B for pt-PT.
Original Dataset: https://huggingface.co/datasets/ybisk/piqa
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/piqa-mt-pt.P3B3
P3B3
Portuguese multi-turn conversational benchmark for measuring European and Brazilian Portuguese variety bias in LLMs.
For more details, see the P3B3 paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/P3B3.xstest_ptpt
XSTest-PT
Portuguese machine translation of XSTest, a benchmark for identifying exaggerated safety behaviors in language models.
Translated using a Finetuned GemmaX2-9B for pt-PT.
Original Dataset: https://huggingface.co/datasets/Paul/XSTest
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/xstest_ptpt.
