datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
P3B3
P3B3
Portuguese multi-turn conversational benchmark for measuring European and Brazilian Portuguese variety bias in LLMs.
For more details, see the P3B3 paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/P3B3.humaneval_mt_pt
HumanEval-PT
Portuguese version of HumanEval, a code generation benchmark with programming problems and test cases.
Translated using NLLB-200.
Original Dataset: https://huggingface.co/datasets/openai/openai_humaneval
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/humaneval_mt_pt.sst_pt-pt
Simple Safety Tests-PT
Portuguese machine translation of Simple Safety Tests, a benchmark for evaluating model safety and harmful content detection.
Translated using a Finetuned GemmaX2-9B for pt-PT.
Original Dataset: https://huggingface.co/datasets/Bertievidgen/SimpleSafetyTests
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/sst_pt-pt.
