datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
portuguese-male-voice-A-datasetreasoning-v1-20m-portugueseglaiveai/reasoning-v1-20m translated to portuguese.
portuguese_benchmark
Portuguese Benchmark
This a collection of datasets in Portuguese initially meant to train and evaluate supervised language models such as BERT, RoBERTa, etc...
It contains 10 datasets and 18 Tasks for Classification (CLS), NLI, Semantic Similarity Scoring (STS) and Named-Entity Recognition (NER).
NER
Classification
NLI
STS
LeNER-Br
HateBR_offensive_binary
assin2-rte
assin2-sts
UlyssesNER-Br-PL-coarse
HateBR_offensive_level
UlyssesNER-Br-C-coarse… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/portuguese_benchmark.BLUEX
BLUEX
There is a repository with the minimal code for using this dataset available here. If you use this dataset for research, please cite the paper:
@misc{almeida2023bluex,
title={BLUEX: A benchmark based on Brazilian Leading Universities Entrance eXams},
author={Thales Sales Almeida and Thiago Laitz and Giovana K. Bonás and Rodrigo Nogueira},
year={2023},
eprint={2307.05410},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
math_dataset_portugueseTo run generation code within 'mathematics_dataset\mathematics_dataset':
Activate python venv .\.venv\Scripts\activate
Requirements defined in requires.txt
Run python generate_to_file.py --output_dir ds to generate dataset to directory \ds
Had to change enconding when opening files to utf-8 so that some characters are allowed (ã õ é)
To obtain dataset with the correct amount of rows:
python generate_to_file.py --output_dir ds --per_train_module 1999998 --per_test_module 10000
This… See the full description on the dataset page: https://huggingface.co/datasets/liaad/math_dataset_portuguese.portuguese-unified-pronunciation-lexicon
Portuguese Unified Pronunciation Lexicon
A flat, single-row-per-pronunciation dataset merging Portuguese IPA transcriptions from three authoritative sources. Each row is a word × region × POS tuple with both broad phonemic (ipa_broad) and narrow phonetic (ipa_narrow) transcriptions normalized across sources.
Source
Words
Convention
Description
Infopédia (Porto Editora)
102,685
Broad phonemic
European Portuguese dictionary IPA
Wiktionary (pt.wiktionary.org)
15,720… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-unified-pronunciation-lexicon.cml_tts_dataset_portugueseSLR-Bench-Portuguese
🧠 SLR-Bench-Portuguese: Scalable Logical Reasoning Benchmark (Portuguese Edition)
SLR-Bench Multilingual Versions:
SLR-Bench-Portuguese is the Portuguese-language pendant of the original SLR-Bench dataset.
It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into Portuguese.
This enables systematic evaluation and training of Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-Portuguese.SPEEED_s3_words_portuguese_0k-90kportuguese_tedx_alignedalpaca-portuguese-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-portuguese-cleaned.raw_dataset_with_embeddings_bert-base-portuguese-cased-nli-assin-2
Dataset Card for "raw_dataset_with_embeddings_bert-base-portuguese-cased-nli-assin-2"
More Information needed
story_cloze_pt
Dataset Card for "story_cloze_pt"
This is a portuguese translation of the xstory_cloze dataset. The translation was performed using the Google Translate API.
This dataset follows the same structure as the original.
portuguese-classic-books-adapted-to-modern-portuguese-brOkay, here is the improved and expanded text translated into American English, including the corrected citation format.
Classic Portuguese Language Books Adapted to Modern Brazilian Portuguese
Detailed Dataset Description
This dataset presents a unique collection of texts derived from classic books of Portuguese language literature, with a strong representation of Brazilian authors. All selected works are in the public domain and were originally sourced from… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/portuguese-classic-books-adapted-to-modern-portuguese-br.portuguese_sentiment_analysisThis dataset is based on the dataset originally posted in Kaggle
portuguese-electionsportuguese_vidportuguese-eval-logs-olmo2-smollm3
Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3
These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs:
SmolLM3
OLMo-2-0425-1B
OLMo-2-1124-7B
Splits
Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.portuguese-dialects-ipa-synthetic
portuguese-dialects-ipa-synthetic
920 dialect-register sentences — 20 per lect across 46 lects: 41 Portuguese varieties
(European regional, insular, Brazilian regional, African/Asian/border national norms,
medieval stages) plus the other languages of Portugal and their kin: Mirandese (3 lects),
Asturleonese of Portugal (Rionorese, Guadramilese) with Barranquenho,
and Galician-Portuguese. Each row carries two IPA columns with distinct provenance.
Schema
sentence —… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-dialects-ipa-synthetic.hatecheck-portuguese
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-portuguese.ticuna-spanish-portuguese
Ticuna (tca) – Spanish – Portuguese Corpus
First text corpus for Ticuna (ISO 639-3 tca), a tonal language
isolate of the Brazil/Colombia/Peru tri-border.
Configs
| Config | Rows |
| parallel | train 43,248 / validation 596 / test 3,238 |
| monolingual | train 46,545 / validation 298 / test 1,613 |
| lexicon | train 10,419 / validation 568 / test 539 |
| instructions | train 52,836 |
| backtranslation | train 33,944 |
The short version of what matters… See the full description on the dataset page: https://huggingface.co/datasets/aimeri/ticuna-spanish-portuguese.portuguese-language-identification-rawEmakhuwa-Portuguese-News-MT
News Parallel Dataset for Emakhuwa of Mozambique
This repository contains releases of parallel data for machine translation in Mozambican languages.
Currently, it supports one language pair, Portuguese-Emakhuwa, Emakhuwa being the widely spoken language in Mozambique.
Dataset Details
Dataset Description
Funded by: This dataset was created with support from Lacuna Fund, the world’s first collaborative effort to provide data scientists, researchers, and… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-News-MT.portuguese-legal-sentences-v0
Work developed as part of Project IRIS.
Thesis: A Semantic Search System for Supremo Tribunal de Justiça
Portuguese Legal Sentences
Collection of Legal Sentences from the Portuguese Supreme Court of Justice
The goal of this dataset was to be used for MLM and TSDAE
Contributions
@rufimelo99
If you use this work, please cite:
@InProceedings{MeloSemantic,
author="Melo, Rui
and Santos, Pedro A.
and Dias, Jo{\~a}o",
editor="Moniz, Nuno
and Vale, Zita
and… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/portuguese-legal-sentences-v0.portuguesechat
Dataset Card for Portuguese Chat
We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved.
Dedicated towards addressing this problem, I release 3 new datasets rishiraj/portuguesechat, rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/portuguesechat.portuguese-edu-qwen-annotations
Annotations for the Portuguese-Edu classifier 📚
Dataset Summary
This dataset contains the annotations used for training an educational classifier (Polygl0t/portuguese-bertimbau-large-edu-classifier and Polygl0t/portuguese-bertimbau-edu-classifier). These annotations were generated by Qwen/Qwen2.5-32B-Instruct.
Supported Tasks and Leaderboards
This dataset can be used for the task of text classification, specifically for educational quality assessment in… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-edu-qwen-annotations.alpaca-gpt4-portugueseThe dataset is used in the research related to MultilingualSIFT.
portuguese-blogs
Dataset Details
Blog-1 may include other languages in an unstructured text format without markdown. The latest one, Blog-6, is formatted in markdown and may contain less other languages text.
Texts are separated by the string <|endoftext|>.
Uses
Training language models.
Dataset Structure
A simple text file with articles separated by <|endoftext|> between each text.
Dataset Creation
First semester of 2024.
Bias, Risks, and Limitations… See the full description on the dataset page: https://huggingface.co/datasets/fabiovilao/portuguese-blogs.portuguese-toxicity-qwen-annotations
Annotations for the Portuguese-Toxicity classifier 📚
Dataset Summary
This dataset contains the annotations used for training a toxicity classifier (Polygl0t/portuguese-bertabaporu-large-toxicity-classifier and Polygl0t/portuguese-bertimbau-toxicity-classifier). These annotations were generated by Qwen/Qwen2.5-32B-Instruct.
Supported Tasks and Leaderboards
This dataset can be used for the task of text classification, specifically for toxicity detection in… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-toxicity-qwen-annotations.PortugueseDollyPortugueseDolly é uma tradição do Databricks Dolly 15k para português brasileiro (pt-br) utilizando o nllb 3.3b.
*Somente para demonstração e pesquisa. Proibido para uso comercial.
PortugueseDolly is a translation of the Databricks Dolly 15k into Brazilian Portuguese (pt-br) using GPT3.5 Turbo.
*For demonstration and research purposes only. Commercial use prohibited.
