amalia-llm/amalia-Dolci-Instruct-SFT
AMALIA Dolci-Instruct-SFT Version of the allenai/Dolci-Instruct-SFT dataset used in the ramp down phase of AMALIA's Supervised Fine-Tuning stage. This dataset was developed by sampling the highest quality entries of the selected splits, and translating part of those entries to European Portuguese. Both the quality classification and translation were done using google/gemma-4-31B-it. This dataset went through a processing pipeline to: Remove entries that reference… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Dolci-Instruct-SFT.
<div align="center">
<img width="400px" src="https://github.com/AMALIA-LLM/amalia-llm.github.io/blob/main/source/_static/logo/logo-color-black.png?raw=true">
    
</div>
AMALIA Dolci-Instruct-SFT
Version of the allenai/Dolci-Instruct-SFT dataset used in the ramp down phase of AMALIA's Supervised Fine-Tuning stage.
This dataset was developed by sampling the highest quality entries of the selected splits, and translating part of those entries to European Portuguese. Both the quality classification and translation were done using google/gemma-4-31B-it.
This dataset went through a processing pipeline to:
- Remove entries that reference other LLMs or research labs;
- Remove entries with tool usage;
Original Dataset: https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT
Note: This dataset comprises machine translated content and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title = "{AMALIA}: A Fully Open Large Language Model for {E}uropean {P}ortuguese",
author = "Simpl{{\'i}}cio, Afonso and Vinagre, Gon{{\c{{c}}}}alo and Ramos, Miguel Moura and Tavares, Diogo and Ferreira, Rafael and Attanasio, Giuseppe and Alves, Duarte M. and Calvo, In{{\^e}}s and Vieira, In{{\^e}}s and Guerra, Rui and Furtado, James and Canaverde, Beatriz and Paulo, Iago and Ramos, Vasco and Gl{{\'o}}ria-Silva, Diogo and Faria, Miguel and Treviso, Marcos and Gomes, Daniel and Gomes, Pedro and Semedo, David and Martins, Andr{{\'e}} and Magalh{{\~a}}es, Jo{{\~a}}o",
booktitle = "Proceedings of the 17th International Conference on Computational Processing of {{P}}ortuguese ({{PROPOR}} 2026) - Vol. 1",
month = apr,
year = "2026",
address = "Salvador, Brazil",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.propor-1.38/",
pages = "380--391",
isbn = "979-8-89176-387-6"
}
@misc{simplicio2026amaliatechnicalreportfully,
title = {AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese},
author = {Afonso Simplício and Gonçalo Vinagre and Miguel Moura Ramos and Diogo Tavares and Rafael Ferreira and Giuseppe Attanasio and Duarte M. Alves and Inês Calvo and Inês Vieira and Rui Guerra and James Furtado and Beatriz Canaverde and Iago Paulo and Vasco Ramos and Diogo Glória-Silva and Miguel Faria and Marcos Treviso and Daniel Gomes and Pedro Gomes and David Semedo and André Martins and João Magalhães},
year = {2026},
eprint = {2603.26511},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2603.26511}
}