datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EGSciQA-ptPT-V1
EGSciQA-ptPT-V1
EGSciQA-ptPT-V1 is the first European Portuguese (pt-PT) supervised fine-tuning
dataset for evidence-grounded scientific question answering and reasoning. Each
example contains a system prompt, an instruction with scientific evidence, and
a structured target response using <raciocinio> and <resposta> sections.
Samples were generated directly from real pt-PT scientific manuscripts available in the amalia-llm/CorEGe-PT corpus.
This repository contains the final… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/EGSciQA-ptPT-V1.hendrycks-math-ptpt
Hendrycks Math MT-PT
Portuguese translated mathematics problems covering algebra, geometry, number theory, and more.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/EleutherAI/hendrycks_math
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/hendrycks-math-ptpt.aerial-solid-waste-dataset-for-classification-and-segmentation
Aerial Solid Waste Dataset for Classification and Segmentation in Cancun, Quintana Roo
size_categories:
- 1K<n<10K
Drones e Inteligencia Artificial para la Detección de Residuos en Zonas Semiurbanas de Cancún
Repositorio oficial del conjunto de datos utilizado para el desarrollo y evaluación de modelos de inteligencia artificial orientados a la detección y localización de residuos sólidos mediante imágenes aéreas capturadas con drones en zonas semiurbanas de… See the full description on the dataset page: https://huggingface.co/datasets/PT-PropuestaModelo-Unicaribe/aerial-solid-waste-dataset-for-classification-and-segmentation.ptpt-failure-set-gate
Where a local 27B actually breaks against a frontier model — a European-Portuguese failure-set gate
On broad everyday tasks, a clean local 27B is near-indistinguishable from a frontier model under blind judging. The gaps that remain are narrow, behavioral, and regex-detectable — which is exactly what small adapters fix. This dataset is the measurement instrument: six hard-sets with deterministic checks, plus the scorer and the methodology write-up.
Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/JigSawPT/ptpt-failure-set-gate.glue-ptptGLUE-PTPT is an European Portuguese translation of the GLUE benchmark using DeepL Pro.MedCalc-Bench-v1.0qa-ptpt
Dataset Card for Dataset Name
Portuguese preprocessed split from MQA dataset containing only the question_title and answer_text columns of records in the ".pt" domain.
The dataset was derived by filtering the following dataset: ju-resplande/qa-pt
The rationale is to have a dataset that is closer aligned with the Portuguese (Portugal) language.
Semantic deduplication splits included for thresholds of 0.7, 0.8 and 0.9 with model… See the full description on the dataset page: https://huggingface.co/datasets/marquesafonso/qa-ptpt.wildguardmix-ptpt
WildGuardTest-PT
Portuguese machine translation of WildGuardTest, a benchmark for evaluating safety guardrails in language models.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/walledai/WildGuardTest
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wildguardmix-ptpt.ptp-nemotron-dense-20260906xstest_ptpt
XSTest-PT
Portuguese machine translation of XSTest, a benchmark for identifying exaggerated safety behaviors in language models.
Translated using a Finetuned GemmaX2-9B for pt-PT.
Original Dataset: https://huggingface.co/datasets/Paul/XSTest
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/xstest_ptpt.FairytaleQA-translated-ptPT
Dataset Card for FairytaleQA-translated-ptPT
Dataset Summary
This repository contains the European Portuguese (pt-PT) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptPT.harmbench-ptpt
HarmBench-PT
Portuguese machine translation of HarmBench (both Standard and Contextual variants), a benchmark for evaluating harmful behavior generation in language models.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/walledai/HarmBench
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/harmbench-ptpt.ptpt-linguistics-if
Portuguese Linguistics Instructions
Dataset handwritten by a Portuguese linguistics expert, covering topics such as phonetics, orthography, wordplay, idiomatic expressions, and grammar classification, specific to European Portuguese.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or AMALIA in your work, please cite:… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/ptpt-linguistics-if.ptp_datatoxichat-ptpt
ToxicChat-PT
Portuguese machine translation of ToxicChat, a benchmark for detecting toxic content in conversational AI.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/lmsys/toxic-chat
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/toxichat-ptpt.aime25-ptpt
AIME-PT 2025
Portuguese translation of problems from the 2025 American Invitational Mathematics Examination (AIME).
Translated using Gemma-4 31B-It.
Original Dataset: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/aime25-ptpt.pt-parliament-interventionsaime-1983-2024-ptpt
AIME-PT (1983-2024)
Portuguese translation of problems from the American Invitational Mathematics Examination (AIME) spanning 1983-2024.
Translated using Gemma-4 31B-It.
Original Dataset: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/aime-1983-2024-ptpt.legal-ptpwikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-1-OP-False-train-perplexitywikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-2-OP-False-train-perplexityaime-2024-ptpt
AIME-PT 2024
Portuguese translation of problems from the 2024 American Invitational Mathematics Examination (AIME).
Translated using Gemma-4 31B-It.
Original Dataset: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/aime-2024-ptpt.aime22-23math-500-ptpt
Math-500-PT
Portuguese machine translation of Math-500, a benchmark of 500 challenging math problems across various topics.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/HuggingFaceH4/MATH-500
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/math-500-ptpt.wikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-2-OP-True-train-perplexitypt-pi0-training-datagemini-dataset-rasalgethi-ptptadaption-ptpn11-tier1-variant-assessments
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ptpn11_tier1_variant_assessments
This dataset contains research-level assessments of PTPN11 missense variants assigned to 'Tier 1' priority based on strict computational evidence filters. Each entry provides structured interpretations including CADD PHRED scores, AlphaMissense predictions, gnomAD frequencies, and functional domain contexts for the SHP-2 protein. The content focuses on… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-ptpn11-tier1-variant-assessments.ptparlThe PTPARL dataset is a dataset containing 5713 interventions in the Portuguese parliament.pt-pt-lexicon
booteek pt-PT lexicon
A SQLite lexicon of European Portuguese (pt-PT) unigrams, bigrams (PMI-ranked),
phrases (3/4/5-grams), and diagnostic tokens (pt-PT vs pt-BR variants), derived from
the booteek-ai/fineweb2-bagaco2
corpus (mirror of duarteocarmo/fineweb2-bagaco2).
Built for grounding LLM outputs in European Portuguese (not Brazilian) — Haiku,
GPT-4, Gemini all default to pt-BR phrasing when asked for "Portuguese" without aggressive
prompting. This lexicon provides a phrase bank… See the full description on the dataset page: https://huggingface.co/datasets/booteek-ai/pt-pt-lexicon.
