CoolFace
Datasetpublic

amalia-llm/AMALIA-LLM-0626-SFT-Dataset

AMALIA LLM Supervised Finetuning Dataset Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training. Base Data Mix This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table: Dataset Count amalia-llm/persona_math 63,731… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-0626-SFT-Dataset.

sourceHugging Faceupdated 3mo agoView on Hugging Face
2likes68downloads
Dataset Card

<div align="center">

<img width="400px" src="https://github.com/AMALIA-LLM/amalia-llm.github.io/blob/main/source/_static/logo/logo-color-black.png?raw=true">

![AMALIA](https://amaliallm.pt/) ![Paper](https://aclanthology.org/2026.propor-1.38/) ![Data](https://huggingface.co/collections/amalia-llm/amalia-lm-eval/) ![Main Repo](https://github.com/AMALIA-LLM/AMALIA) ![Eval Repo](https://github.com/AMALIA-LLM/amalia-lm-eval)

</div>

AMALIA LLM Supervised Finetuning Dataset

Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training.

Base Data Mix

This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table:

DatasetCount
amalia-llm/persona_math63,731
amalia-llm/persona_nemotron101,885
amalia-llm/persona_general156,397
amalia-llm/personainstructionfollowing9,084
amalia-llm/Amalia_hardcoded780 (upsampled 5x)
amalia-llm/ptpt-linguistics-if200
amalia-llm/PT-Culture_Data26,089
amalia-llm/wikipedia_conversations97,553
amalia-llm/wikipedia-rag58,549
amalia-llm/amalia-PTradutor50,000
amalia-llm/wmt24ppxxtopt10k10,000
amalia-llm/amalia-smoltalk3,507
amalia-llm/amalia-smol-rewrite-PT53,342
amalia-llm/amalia-smolsummarizept96,356
amalia-llm/amalia-smoltalk2everydayconv_pt305
amalia-llm/amalia-smoltalk213,203
amalia-llm/amalia-Nemotron-Instruction-Following-Chat-v193,763
amalia-llm/amalia-Nemotron-SFT-Instruction-Following-Chat-v21,970,977
amalia-llm/amalia-Nemotron-Math-Proofs-v1460,113
amalia-llm/amalia-Nemotron-SFT-Math-v3924,449
amalia-llm/amalia-Nemotron-SFT-Competitive-Programming-v2841,334
amalia-llm/amalia-Nemotron-SFT-SWE-v2204,811
amalia-llm/amalia-Nemotron-SpecializedDomains-Finance-v166,167
amalia-llm/amalia-Nemotron-SFT-Multilingual-v11,029,778
amalia-llm/amalia-Nemotron-SFT-Safety-v144,131
amalia-llm/amalia-Nemotron-Science-v1226,319

Ramp Down Data Mix

This data mix comprises samples of datasets used in the base data mix, whether using the complete datasets, a random sample or selecting the highest quality entries. To classify the quality of the entries, the model google/gemma-4-31B-it was used as a quality judge. A part of the high quality entries was translated to European portuguese also using google/gemma-4-31B-it. The dataset counts and entry-selection methods are described in the following table:

DatasetCount
amalia-llm/amalia-Dolci-Instruct-SFT196,924 (best)<br>43,074 (translated)
amalia-llm/amalia-Nemotron-SFT-Instruction-Following-Chat-v2126,314 (random)<br>73,645 (translated)
amalia-llm/hermes3specialsystem_prompts55,000 (best)<br>51,194 (translated)
amalia-llm/wikipedia_conversations97,553
amalia-llm/PT-Culture_Data34,505
amalia-llm/amalia-PTradutor20,000
amalia-llm/smol-rewrite-PT10,000
amalia-llm/personainstructionfollowing9,084
amalia-llm/wikipedia-rag8,492
amalia-llm/Amalia_hardcoded780 (upsampled 5x)

Note: This dataset comprises machine translated content and may contain translation errors or artifacts.

This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.


Citation

If you use this dataset or AMALIA in your work, please cite:

bibtex
@inproceedings{simplicio-etal-2026-amalia,
    title = "{AMALIA}: A Fully Open Large Language Model for {E}uropean {P}ortuguese",
    author = "Simpl{{\'i}}cio, Afonso and Vinagre, Gon{{\c{{c}}}}alo and Ramos, Miguel Moura and Tavares, Diogo and Ferreira, Rafael and Attanasio, Giuseppe and Alves, Duarte M. and Calvo, In{{\^e}}s and Vieira, In{{\^e}}s and Guerra, Rui and Furtado, James and Canaverde, Beatriz and Paulo, Iago and Ramos, Vasco and Gl{{\'o}}ria-Silva, Diogo and Faria, Miguel and Treviso, Marcos and Gomes, Daniel and Gomes, Pedro and Semedo, David and Martins, Andr{{\'e}} and Magalh{{\~a}}es, Jo{{\~a}}o",
    booktitle = "Proceedings of the 17th International Conference on Computational Processing of {{P}}ortuguese ({{PROPOR}} 2026) - Vol. 1",
    month = apr,
    year = "2026",
    address = "Salvador, Brazil",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.propor-1.38/",
    pages = "380--391",
    isbn = "979-8-89176-387-6"
}

@misc{simplicio2026amaliatechnicalreportfully,
    title = {AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese},
    author = {Afonso Simplício and Gonçalo Vinagre and Miguel Moura Ramos and Diogo Tavares and Rafael Ferreira and Giuseppe Attanasio and Duarte M. Alves and Inês Calvo and Inês Vieira and Rui Guerra and James Furtado and Beatriz Canaverde and Iago Paulo and Vasco Ramos and Diogo Glória-Silva and Miguel Faria and Marcos Treviso and Daniel Gomes and Pedro Gomes and David Semedo and André Martins and João Magalhães},
    year = {2026},
    eprint = {2603.26511},
    archivePrefix = {arXiv},
    primaryClass = {cs.CL},
    url = {https://arxiv.org/abs/2603.26511}
}