datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
portuguese_benchmark
Portuguese Benchmark
This a collection of datasets in Portuguese initially meant to train and evaluate supervised language models such as BERT, RoBERTa, etc...
It contains 10 datasets and 18 Tasks for Classification (CLS), NLI, Semantic Similarity Scoring (STS) and Named-Entity Recognition (NER).
NER
Classification
NLI
STS
LeNER-Br
HateBR_offensive_binary
assin2-rte
assin2-sts
UlyssesNER-Br-PL-coarse
HateBR_offensive_level
UlyssesNER-Br-C-coarse… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/portuguese_benchmark.SLR-Bench-Portuguese
🧠 SLR-Bench-Portuguese: Scalable Logical Reasoning Benchmark (Portuguese Edition)
SLR-Bench Multilingual Versions:
SLR-Bench-Portuguese is the Portuguese-language pendant of the original SLR-Bench dataset.
It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into Portuguese.
This enables systematic evaluation and training of Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-Portuguese.portuguese-eval-logs-olmo2-smollm3
Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3
These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs:
SmolLM3
OLMo-2-0425-1B
OLMo-2-1124-7B
Splits
Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.hatecheck-portuguese
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-portuguese.ticuna-spanish-portuguese
Ticuna (tca) – Spanish – Portuguese Corpus
First text corpus for Ticuna (ISO 639-3 tca), a tonal language
isolate of the Brazil/Colombia/Peru tri-border.
Configs
| Config | Rows |
| parallel | train 43,248 / validation 596 / test 3,238 |
| monolingual | train 46,545 / validation 298 / test 1,613 |
| lexicon | train 10,419 / validation 568 / test 539 |
| instructions | train 52,836 |
| backtranslation | train 33,944 |
The short version of what matters… See the full description on the dataset page: https://huggingface.co/datasets/aimeri/ticuna-spanish-portuguese.Emakhuwa-Portuguese-News-MT
News Parallel Dataset for Emakhuwa of Mozambique
This repository contains releases of parallel data for machine translation in Mozambican languages.
Currently, it supports one language pair, Portuguese-Emakhuwa, Emakhuwa being the widely spoken language in Mozambique.
Dataset Details
Dataset Description
Funded by: This dataset was created with support from Lacuna Fund, the world’s first collaborative effort to provide data scientists, researchers, and… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-News-MT.portuguese-tweetstest_split_with_embeddings_bert_base_portuguese
Dataset Card for "test_split_with_embeddings_bert_base_portuguese"
More Information needed
PortugueseMMLU
Dataset Components
The dataset is partitioned into three discrete tables stored in CSV or Parquet format:
Questions
Recipes
Evaluation Results
Each component is described in detail below.
Questions
area
domain
question_number
An integer index uniquely identifying each question inside the knowledge domain.
translation_method
English, Google Translate, GPT-3.5-Turbo, GPT-4o, Human
question
option_a, option_b, option_c, option_d
Recipes
area… See the full description on the dataset page: https://huggingface.co/datasets/TaigoPedrosa/PortugueseMMLU.portuguese-instruct-quality-qwen-annotations
Annotations for the Portuguese-instruct-quality classifier 📚
Dataset Summary
This dataset contains the annotations used for training a quality filter for instruction type data (Polygl0t/portuguese-qwen3-4b-instruct-quality-classifier and Polygl0t/portuguese-qwen3-4b-instruct-quality-judge).
These annotations were generated by Qwen/Qwen2.5-32B-Instruct.
Supported Tasks and Leaderboards
This dataset can be used for the task of text classification, or for… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-instruct-quality-qwen-annotations.SLR-Bench-Portuguese
🧠 SLR-Bench-Portuguese: Scalable Logical Reasoning Benchmark (Portuguese Edition)
SLR-Bench Versions:
SLR-Bench-Portuguese is the Portuguese-language pendant of the original SLR-Bench dataset.
It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into Portuguese.
This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/ahmad21omar/SLR-Bench-Portuguese.orpheus-tts-portuguese-datasetportuguese-hate-speech-superset
Portuguese Hate Speech Superset
This dataset is a superset (N=43,222) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Portuguese hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/portuguese-hate-speech-superset.xpaws_pt
Dataset Card for "xpaws_pt"
This is a portuguese translation of the x-paws dataset. The translation was performed using the Google Translate API.
This dataset follows the same structure as the original.
fineweb-portuguese-100k
FineWeb2 Portuguese 100k - Safety Classified
A 100,000-sample subset of FineWeb2 Portuguese web text, classified for content safety using Cohere Command A.
Dataset Description
Each record contains the original FineWeb2 text and metadata, plus a classification field with:
Field
Description
safety_rating
"safe" or "unsafe"
category
List of applicable harm categories (null if safe)
reason
Brief explanation of the classification
Safety Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/fineweb-portuguese-100k.crime_tweets_in_portuguese
DataCrimeBR: Building a Dataset of Crimes Reported in Tweets in Brazil
This dataset contains 61.715 tweets related to possible crime reports, labeled with categories such as "Assalto", "Roubo", "Furto", "Assédio", "Segurança Pública", "Homicídio, and "Outros", along with sentiment analysis, toxicity analysis, and location identification.
A particular feature in the Portuguese language is that many words potentially related to crimes are used in non-criminal contexts, such as "O… See the full description on the dataset page: https://huggingface.co/datasets/miguelribeirokk/crime_tweets_in_portuguese.Nemotron-Safety-Guard-Dataset-v3-portuguese
Nemotron Portuguese Safety (Translated)
Portuguese safety prompts/responses (translated from Spanish), with labels and categories.
Dataset Description
nemotron_pt
Each record includes Portuguese prompt/response text plus safety labels/categories.
Field
Description
id
Example id
prompt
Portuguese prompt text
response
Portuguese response text (may be null)
prompt_label
"safe" or "unsafe"
response_label
"safe" or "unsafe" (may be empty if… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/Nemotron-Safety-Guard-Dataset-v3-portuguese.brazilian-portuguese-anaphylaxis
Brazilian Portuguese Clinical Notes for Anaphylaxis Detection
A dataset of 969 Brazilian Portuguese clinical narratives annotated for the presence or absence of anaphylaxis, built for research in clinical Natural Language Processing (NLP).
Overview
Anaphylaxis is an acute, potentially life-threatening allergic reaction that requires rapid recognition in clinical settings. Automatic detection of anaphylaxis in clinical narratives can support large-scale analysis of… See the full description on the dataset page: https://huggingface.co/datasets/matos1012/brazilian-portuguese-anaphylaxis.fineweb2-portuguese-safetyportuguese-parliament-interventionswildguardtest-portuguesealmeida_portuguese_pt
Almeida Atualizada (Portuguese)
Description
A Almeida Atualizada is a revised edition of the classic Portuguese Bible translation by João Ferreira de Almeida (1628-1691), a Portuguese Protestant pastor. Almeida's original translation, the first complete Portuguese Bible, was published in 1748-1753. The Almeida Atualizada (updated Almeida) modernizes the language while preserving the fidelity to the original Hebrew and Greek texts. It is the most widely used… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/almeida_portuguese_pt.fineweb-edu-gemini-annotations-portuguese-regressionfineweb-portuguese-chunkportuguese-hate-speechNemotron-Safety-Guard-Dataset-v3-portuguese-safetyzeroshot_portugueseportuguese_hate_speech_lighteval_fewshotmls-portuguese-snacportuguese-hate-speech
