CoolFace
Datasetpublic

amalia-llm/smol-rewrite-PT

SMOL Rewrite PT This dataset is the translated version of the smol-rewrite subset of the HuggingFaceTB/smoltalk. This dataset includes an high-quality split used in the ramp down phase of the AMALIA's model post-training. The quality classification was done using google/gemma-4-31B-it. Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk Note: This dataset comprises machine translated content and may contain translation errors or artifacts. This… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol-rewrite-PT.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes28downloads
Dataset Card

<div align="center">

<img width="400px" src="https://github.com/AMALIA-LLM/amalia-llm.github.io/blob/main/source/_static/logo/logo-color-black.png?raw=true">

![AMALIA](https://amaliallm.pt/) ![Paper](https://aclanthology.org/2026.propor-1.38/) ![Data](https://huggingface.co/collections/amalia-llm/amalia-lm-eval/) ![Main Repo](https://github.com/AMALIA-LLM/AMALIA) ![Eval Repo](https://github.com/AMALIA-LLM/amalia-lm-eval)

</div>

SMOL Rewrite PT

This dataset is the translated version of the smol-rewrite subset of the HuggingFaceTB/smoltalk.

This dataset includes an high-quality split used in the ramp down phase of the AMALIA's model post-training. The quality classification was done using google/gemma-4-31B-it.

Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk

Note: This dataset comprises machine translated content and may contain translation errors or artifacts.

This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.


Citation

If you use this dataset or AMALIA in your work, please cite:

bibtex
@inproceedings{simplicio-etal-2026-amalia,
    title = "{AMALIA}: A Fully Open Large Language Model for {E}uropean {P}ortuguese",
    author = "Simpl{{\'i}}cio, Afonso and Vinagre, Gon{{\c{{c}}}}alo and Ramos, Miguel Moura and Tavares, Diogo and Ferreira, Rafael and Attanasio, Giuseppe and Alves, Duarte M. and Calvo, In{{\^e}}s and Vieira, In{{\^e}}s and Guerra, Rui and Furtado, James and Canaverde, Beatriz and Paulo, Iago and Ramos, Vasco and Gl{{\'o}}ria-Silva, Diogo and Faria, Miguel and Treviso, Marcos and Gomes, Daniel and Gomes, Pedro and Semedo, David and Martins, Andr{{\'e}} and Magalh{{\~a}}es, Jo{{\~a}}o",
    booktitle = "Proceedings of the 17th International Conference on Computational Processing of {{P}}ortuguese ({{PROPOR}} 2026) - Vol. 1",
    month = apr,
    year = "2026",
    address = "Salvador, Brazil",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.propor-1.38/",
    pages = "380--391",
    isbn = "979-8-89176-387-6"
}

@misc{simplicio2026amaliatechnicalreportfully,
    title = {AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese},
    author = {Afonso Simplício and Gonçalo Vinagre and Miguel Moura Ramos and Diogo Tavares and Rafael Ferreira and Giuseppe Attanasio and Duarte M. Alves and Inês Calvo and Inês Vieira and Rui Guerra and James Furtado and Beatriz Canaverde and Iago Paulo and Vasco Ramos and Diogo Glória-Silva and Miguel Faria and Marcos Treviso and Daniel Gomes and Pedro Gomes and David Semedo and André Martins and João Magalhães},
    year = {2026},
    eprint = {2603.26511},
    archivePrefix = {arXiv},
    primaryClass = {cs.CL},
    url = {https://arxiv.org/abs/2603.26511}
}