datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reasoning-v1-20m-portugueseglaiveai/reasoning-v1-20m translated to portuguese.
portuguese-unified-pronunciation-lexicon
Portuguese Unified Pronunciation Lexicon
A flat, single-row-per-pronunciation dataset merging Portuguese IPA transcriptions from three authoritative sources. Each row is a word × region × POS tuple with both broad phonemic (ipa_broad) and narrow phonetic (ipa_narrow) transcriptions normalized across sources.
Source
Words
Convention
Description
Infopédia (Porto Editora)
102,685
Broad phonemic
European Portuguese dictionary IPA
Wiktionary (pt.wiktionary.org)
15,720… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-unified-pronunciation-lexicon.portuguese-classic-books-adapted-to-modern-portuguese-brOkay, here is the improved and expanded text translated into American English, including the corrected citation format.
Classic Portuguese Language Books Adapted to Modern Brazilian Portuguese
Detailed Dataset Description
This dataset presents a unique collection of texts derived from classic books of Portuguese language literature, with a strong representation of Brazilian authors. All selected works are in the public domain and were originally sourced from… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/portuguese-classic-books-adapted-to-modern-portuguese-br.portuguese-ocr-datasettask_categories:
image-to-text
task_ids:
optical-character-recognition
text-recognition
Portuguese OCR Dataset
A comprehensive dataset for Portuguese OCR (Optical Character Recognition) generated from classic Portuguese literature with diverse fonts and visual styles.
Dataset Description
This dataset contains 20000 text images for OCR training, created from Portuguese books from Project Gutenberg. Each image contains a complete Portuguese sentence with proper… See the full description on the dataset page: https://huggingface.co/datasets/mazafard/portuguese-ocr-dataset.portuguesechat
Dataset Card for Portuguese Chat
We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved.
Dedicated towards addressing this problem, I release 3 new datasets rishiraj/portuguesechat, rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/portuguesechat.portuguese-blogs
Dataset Details
Blog-1 may include other languages in an unstructured text format without markdown. The latest one, Blog-6, is formatted in markdown and may contain less other languages text.
Texts are separated by the string <|endoftext|>.
Uses
Training language models.
Dataset Structure
A simple text file with articles separated by <|endoftext|> between each text.
Dataset Creation
First semester of 2024.
Bias, Risks, and Limitations… See the full description on the dataset page: https://huggingface.co/datasets/fabiovilao/portuguese-blogs.orca-math-portuguese-64ktranslated for:
Repository: microsoft/orca-math-word-problems-200k
Paper: Orca-Math: Unlocking the potential of
SLMs in Grade School Math
mirror-rhaymison__orca-math-portuguese-64ktranslated for:
Repository: microsoft/orca-math-word-problems-200k
Paper: Orca-Math: Unlocking the potential of
SLMs in Grade School Math
portuguese-qa-instruct-500
Portuguese Q&A Instruction Dataset (500 pairs)
500 Portuguese (PT-PT) question-answer pairs formatted for instruction fine-tuning of language models.
Dataset Structure
Each example has three columns:
Column
Description
Example
instruction
The question in Portuguese
"Qual e a capital de Portugal?"
response
The answer in Portuguese
"A capital de Portugal e Lisboa."
text
Pre-formatted instruction template (see below)
"<|im_start|>user\n..."… See the full description on the dataset page: https://huggingface.co/datasets/nelsondiasandre/portuguese-qa-instruct-500.PortugueseMMLU
Dataset Components
The dataset is partitioned into three discrete tables stored in CSV or Parquet format:
Questions
Recipes
Evaluation Results
Each component is described in detail below.
Questions
area
domain
question_number
An integer index uniquely identifying each question inside the knowledge domain.
translation_method
English, Google Translate, GPT-3.5-Turbo, GPT-4o, Human
question
option_a, option_b, option_c, option_d
Recipes
area… See the full description on the dataset page: https://huggingface.co/datasets/TaigoPedrosa/PortugueseMMLU.portuguese_phonetic_lexicon
📚 Portuguese Phonetic Lexicon Dataset
This dataset contains phonetic and morphological information for Portuguese words, collected from the Portal da Língua Portuguesa. It was generated by scraping the site across multiple Portuguese-speaking regions and dialects.
🌍 Regional Coverage
The dataset includes words as spoken in ten regional variants:
🇵🇹 Lisbon (Standard and Non-Standard)
🇦🇴 Luanda
🇧🇷 Rio de Janeiro (Standard and Non-Standard)
🇧🇷 São Paulo (Standard… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese_phonetic_lexicon.Emakhuwa-Portuguese-OCR-post-correctionBibTeX:
The dataset paper was published in EMNLP 2024.
Please cite as:
@inproceedings{ali-etal-2024-building,
title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks",
author = "Ali, Felermino D. M. A. and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-OCR-post-correction.realtor-conversational-portuguese_br
Realtor Conversational (Portuguese BR) - Realtor-Client Conversation Dataset
Detailed Dataset Description
Introduction:
This dataset, named "Realtor Conversational (Portuguese BR)", offers rich and detailed simulations of conversational interactions between real estate agents (realtors) and clients in Brazil. Generated using the advanced language model gpt-4o-mini, the data is synthetic but designed to mirror the dynamics, vocabulary, and common scenarios found in the… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/realtor-conversational-portuguese_br.Reasoning-MiniGPT-brazilian-portugueserealtor-conversational-portuguese_br
Realtor Conversational (Portuguese BR) - Realtor-Client Conversation Dataset
Detailed Dataset Description
Introduction:
This dataset, named "Realtor Conversational (Portuguese BR)", offers rich and detailed simulations of conversational interactions between real estate agents (realtors) and clients in Brazil. Generated using the advanced language model gpt-4o-mini, the data is synthetic but designed to mirror the dynamics, vocabulary, and common scenarios found in the… See the full description on the dataset page: https://huggingface.co/datasets/rishabmishrasensation/realtor-conversational-portuguese_br.portuguese-sentences-synthetic-g2p
Dataset Card for 'TigreGotico/portuguese_g2p'
Dataset Description
Dataset Summary
TigreGotico/portuguese_g2p is a Grapheme-to-Phoneme (G2P) dataset for Portuguese, offering phonetic transcriptions for sentences across ten different regional variants.
It is derived from the portuguese_phonetic_lexicon and is designed to aid in the development of robust Speech Recognition (ASR) and Text-to-Speech (TTS) models that account for dialectal variation in Portuguese.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-sentences-synthetic-g2p.
