datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bigbenchhard-mt-pt
BBH-PT (Big-Bench Hard)
Portuguese machine translation of BIG-Bench Hard, a challenging subset of the BIG-Bench benchmark covering diverse reasoning tasks.
Translated using a Finetuned GemmaX2-9B for pt-PT with rule-based adaptations.
Note: Some tasks (e.g., hyperbaton) are not translated as they do not transfer meaningfully to Portuguese.
Original Dataset: https://github.com/suzgunmirac/BIG-Bench-Hard
Note: This dataset is machine translated and may contain… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/bigbenchhard-mt-pt.pt_exams
PHEB - Portuguese High School Exams MCQ
MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum.
For more details, see the PHEB paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.alba_mcq
ALBA MCQ
Multiple Choice Version of the ALBA benchmark, a Portuguese language benchmark for proficiency in pt-PT linguistic-related tasks.
For more details, see the ALBA paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/alba_mcq.AMALIA-LLM-0626-SFT-Dataset
AMALIA LLM Supervised Finetuning Dataset
Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training.
Base Data Mix
This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table:
Dataset
Count
amalia-llm/persona_math
63,731… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-0626-SFT-Dataset.persona_nemotron
Persona Nemotron PT Datasets
This is a collection of Portuguese synthetic datasets, consisting of 3 datasets, one with general questions from varied topics, one with math questions, and one with instruction-following requests.
The prompts were generated using an approach similar to PersonaHub, with a translated version of Nemotron Personas. Both prompts and answers were generated using Gemma 3-27B.
This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_nemotron.PT-Culture_Data
Portuguese Cultural SFT Dataset
A supervised fine-tuning dataset for European Portuguese (PT-PT) cultural knowledge, built to teach models the traditions, figures, places, and expressions of Portuguese culture.
The dataset contains 216,832 examples across two versions in conversational SFT format, organised into ten cultural domains:
Domain
v1 (95,815)
v2 (121,017)
Total
Personalities
36,370
62,788
99,158
Audiovisual
15,786
26,857
42,643
Heritage
8,020… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/PT-Culture_Data.smoltalk2_everyday_conv_pt
SMOL Everyday Conversation PT
This dataset consists of a Portuguese version of the smoltalk_smollm3_everyday_conversations_no_think split of HuggingFaceTB/smoltalk2. The first two user turns were translated as well as the first assistant turn, the continuation of the conversation was generated using Gemma 3-27B.
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smoltalk2_everyday_conv_pt.wikipedia-rag
Wikipedia RAG
This dataset consists of raw wikipedia pages with one question and answer for each page. The answers were generated using Gemma 3-27B.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title = "{AMALIA}: A Fully Open Large Language Model for… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wikipedia-rag.saudade-pt
SAUDADE
Portuguese benchmark for temporal reasoning and understanding of temporal relationships across Portuguese events.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title = "{AMALIA}: A Fully Open… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/saudade-pt.amalia-Dolci-Instruct-SFT
AMALIA Dolci-Instruct-SFT
Version of the allenai/Dolci-Instruct-SFT dataset used in the ramp down phase of AMALIA's Supervised Fine-Tuning stage.
This dataset was developed by sampling the highest quality entries of the selected splits, and translating part of those entries to European Portuguese. Both the quality classification and translation were done using google/gemma-4-31B-it.
This dataset went through a processing pipeline to:
Remove entries that reference… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Dolci-Instruct-SFT.persona_instruction_following
Persona Instruction Following Datasets
This is a synthetic instruction-following dataset, available in two configs: full and filtered. Each config contains two language splits, English (en) and Portuguese (pt). The filtered version keeps only the higher-quality examples (quality score 5).
The prompts were generated using an approach similar to PersonaHub, with a translated version of proj-persona/PersonaHub. Both prompts and answers were generated using Gemma 3-27B.… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_instruction_following.amalia-Nemotron-SFT-Math-v3
AMALIA Nemotron-SFT-Math-v3
Version of the nvidia/Nemotron-SFT-Math-v3 dataset used in the AMALIA's Supervised Fine-Tuning stage.
This dataset went through a processing pipeline to:
Remove entries with tool usage;
Remove entries that reference other LLMs or research labs;
Removed the reasoning_content field;
Drop extra unnecessary columns;
Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Math-v3
This dataset is provided as part of the AMALIA… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SFT-Math-v3.DPO-Dataset
AMALIA DPO Dataset
This is the DPO (preference optimization) dataset used to train AMALIA-DPO.
It is a mix of preference pairs, mainly in European Portuguese and English, covering general conversation, instruction following, math, and safety. These pairs come from different sources, including prompts from the SFT mix, responses generated by different models, including an early version of the model, and some public datasets.
This dataset is provided as part of the AMALIA project… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/DPO-Dataset.amalia-Nemotron-Instruction-Following-Chat-v1
AMALIA Nemotron-Instruction-Following-Chat-v1
Version of the nvidia/Nemotron-Instruction-Following-Chat-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage.
This dataset went through a processing pipeline to:
Remove entries where the capability_target field was 'chat' from the chat_if split;
Remove entries that reference other LLMs or research labs;
Remove entries that contained the string '/imagine prompt:';
Removed the reasoning_content field;
Original… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-Instruction-Following-Chat-v1.hermes3_special_system_prompts
Hermes 3 SFT Special System Prompts
Version of the NousResearch/Hermes-3-Dataset dataset used in AMALIA's Supervised Fine-Tuning stage.
This subset was manually selected in order to retain only entries containing custom system prompts that substantially modify the model’s behavior. It were also built splits with the highest quality entries and a translated part of those entries to European Portuguese. Both the quality classification and translation were done using… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/hermes3_special_system_prompts.piqa-mt-pt
PIQA-PT
Portuguese machine translation of PIQA (Physical Interaction QA), a benchmark for physical commonsense reasoning.
Translated using a Finetuned GemmaX2-9B for pt-PT.
Original Dataset: https://huggingface.co/datasets/ybisk/piqa
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/piqa-mt-pt.amalia-Nemotron-Science-v1
AMALIA Nemotron-Science-v1
Version of the nvidia/Nemotron-Science-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage.
This dataset went through a processing pipeline to:
Remove entries that reference other LLMs or research labs;
Remove the reasoning_content field;
Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-Science-v1
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-Science-v1.amalia-PTradutor
AMALIA PTradutor
This dataset consists of translations from English to Portuguese and from Portuguese to English. These pairs come from liaad/PTradutor and were transformed into a conversation format by adding an instruction to translate the text, placed in either the user message or the system message.
This dataset includes an high-quality split used in the ramp down phase of the AMALIA's model post-training. The quality classification was done using… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-PTradutor.amalia-pilot-honesty-v2
AMALIA pilot — honesty vector datasets (v1 refusals + v2 corrective mix)
Training data from the first two iterations of a verifier-gated fine-tuning
pilot on AMALIA-9B-0626-DPO,
targeting identity/fact confabulation (the model's weakest measured behavior:
43.3% on our honesty harness). Full methodology, harness, and reports:
github.com/teex-pt/pt-amalia.
These are research pilot artifacts — small by design (the pilot validates
the loop, not the scale). Every sample was produced… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-pilot-honesty-v2.Amalia_hardcoded
AMALIA Hardcoded
Handwritten dataset to provide the AMALIA model with self-referential knowledge about its development and capabilities, both in Portuguese and English.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title = "{AMALIA}: A Fully Open Large Language… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/Amalia_hardcoded.AMALIA-LLM-1225-SFT-Dataset
AMALIA LLM 1225 Supervised Finetuning Dataset
Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model version released in December 2025.
This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table:
Dataset
Count
Persona-PT Instruction Following
9,084
Persona-EN Instruction Following
34,704
Persona Nemotron Instruction Following
4,483… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-1225-SFT-Dataset.pt_text_completion
PT-PT Completions
Simple text-completion dataset to evaluete model bias towards European Portuguese (pt-PT) or Brazilian Portuguese (pt-BR).
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title =… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_text_completion.wikipedia_conversations
Wikipedia Conversations
This is a dataset of conversations based on Portuguese Wikipedia articles.
The conversations were generated using Gemma 3-27B.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title = "{AMALIA}: A Fully Open Large Language Model for… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wikipedia_conversations.openbookqa-mt-pt
OpenBookQA-PT
Portuguese machine translation of OpenBookQA, a question-answering dataset modeled after open-book exams for elementary science.
Translated using a Finetuned GemmaX2-9B for pt-PT.
Original Dataset: https://huggingface.co/datasets/allenai/openbookqa
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/openbookqa-mt-pt.PersonaHub
Scaling Synthetic Data Creation with 1,000,000,000 Personas
This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas:
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/AmaliaE/PersonaHub.amalia-Nemotron-SFT-Instruction-Following-Chat-v2
Nemotron-SFT-Instruction-Following-Chat-v2
Version of the nvidia/Nemotron-SFT-Instruction-Following-Chat-v2 dataset used in the AMALIA's Supervised Fine-Tuning stage, both in the base and ramp down stages. The ramp down stage comprised a randomly select subset of the base subset, where part was translated to European Portuguese using google/gemma-4-31B-it.
This dataset went through a processing pipeline to:
Remove entries that reference other LLMs or research labs;… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SFT-Instruction-Following-Chat-v2.smol-rewrite-PT
SMOL Rewrite PT
This dataset is the translated version of the smol-rewrite subset of the HuggingFaceTB/smoltalk.
This dataset includes an high-quality split used in the ramp down phase of the AMALIA's model post-training. The quality classification was done using google/gemma-4-31B-it.
Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
Note: This dataset comprises machine translated content and may contain translation errors or artifacts.
This… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol-rewrite-PT.smol_summarize_pt
SMOL Summarize PT
This dataset is the translated version of the smol-summarize subset of the HuggingFaceTB/smoltalk.
Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
Note: This dataset comprises machine translated content and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol_summarize_pt.amalia-Nemotron-SpecializedDomains-Finance-v1
AMALIA Nemotron-SpecializedDomains-Finance-v1
Version of the nvidia/Nemotron-SpecializedDomains-Finance-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage.
This dataset went through a processing pipeline to:
Remove entries that reference other LLMs or research labs;
Remove the reasoning_content field;
Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1
This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SpecializedDomains-Finance-v1.persona_math
Persona Math
This is a mathematical European Portuguese synthetic dataset. The datasets were generated using an approach similar to PersonaHub, with a translated version of proj-persona/PersonaHub. Both prompts and answers were generated using Gemma 3-27B.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or AMALIA in your work, please cite:… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_math.
