datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wildguardmix-ptpt
WildGuardTest-PT
Portuguese machine translation of WildGuardTest, a benchmark for evaluating safety guardrails in language models.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/walledai/WildGuardTest
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wildguardmix-ptpt.toxichat-ptpt
ToxicChat-PT
Portuguese machine translation of ToxicChat, a benchmark for detecting toxic content in conversational AI.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/lmsys/toxic-chat
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/toxichat-ptpt.aime-1983-2024-ptpt
AIME-PT (1983-2024)
Portuguese translation of problems from the American Invitational Mathematics Examination (AIME) spanning 1983-2024.
Translated using Gemma-4 31B-It.
Original Dataset: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/aime-1983-2024-ptpt.
