datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AMALIA-VL-SFT-Dataset
AMALIA-VL-Training-Dataset
Dataset Description
This dataset is provided as part of the AMALIA project.
This is the vision+language training mix for AMALIA-VL-SFT. Each subset is one
source dataset in the mix, each with a single train split. The only datasets that are absent from this mix are those that derive directly from the core LLM training mix, and can be found in the AMALIA-LLM Post Training Collection.
Example usage:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-VL-SFT-Dataset.pt_exams
PHEB - Portuguese High School Exams MCQ
MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum.
For more details, see the PHEB paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.alba_mcq
ALBA MCQ
Multiple Choice Version of the ALBA benchmark, a Portuguese language benchmark for proficiency in pt-PT linguistic-related tasks.
For more details, see the ALBA paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/alba_mcq.wildguardmix-ptpt
WildGuardTest-PT
Portuguese machine translation of WildGuardTest, a benchmark for evaluating safety guardrails in language models.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/walledai/WildGuardTest
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wildguardmix-ptpt.DPO-Dataset
AMALIA DPO Dataset
This is the DPO (preference optimization) dataset used to train AMALIA-DPO.
It is a mix of preference pairs, mainly in European Portuguese and English, covering general conversation, instruction following, math, and safety. These pairs come from different sources, including prompts from the SFT mix, responses generated by different models, including an early version of the model, and some public datasets.
This dataset is provided as part of the AMALIA project… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/DPO-Dataset.openbookqa-mt-pt
OpenBookQA-PT
Portuguese machine translation of OpenBookQA, a question-answering dataset modeled after open-book exams for elementary science.
Translated using a Finetuned GemmaX2-9B for pt-PT.
Original Dataset: https://huggingface.co/datasets/allenai/openbookqa
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/openbookqa-mt-pt.toxichat-ptpt
ToxicChat-PT
Portuguese machine translation of ToxicChat, a benchmark for detecting toxic content in conversational AI.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/lmsys/toxic-chat
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/toxichat-ptpt.aime-1983-2024-ptpt
AIME-PT (1983-2024)
Portuguese translation of problems from the American Invitational Mathematics Examination (AIME) spanning 1983-2024.
Translated using Gemma-4 31B-It.
Original Dataset: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/aime-1983-2024-ptpt.amalia-smoltalk
AMALIA smoltalk
Version of the smol-magpie-ultra subset of HuggingFaceTB/smoltalk dataset used in the AMALIA's Supervised Fine-Tuning stage.
This dataset went through a processing pipeline to keep only the entries labeled with excellent quality;
Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-smoltalk.amalia-cita-legal
AMALIA cita-legal — grounded legal citation SFT (RAG-first)
(question + real source excerpts → answer that cites the exact article, or a refusal when the excerpts don't answer the question). Built for
specializing AMALIA-9B
toward Portuguese legal text, as the grounded-answering counterpart to
teex-pt/amalia-sum-dre.
Derived from
teex-pt/leis-pt-consolidada
by teex-pt/pt-amalia.
Why this exists (RAG-first, not closed-book)
leis-pt's own project spec concludes that… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-cita-legal.
