datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
portuguese-unified-pronunciation-lexicon
Portuguese Unified Pronunciation Lexicon
A flat, single-row-per-pronunciation dataset merging Portuguese IPA transcriptions from three authoritative sources. Each row is a word × region × POS tuple with both broad phonemic (ipa_broad) and narrow phonetic (ipa_narrow) transcriptions normalized across sources.
Source
Words
Convention
Description
Infopédia (Porto Editora)
102,685
Broad phonemic
European Portuguese dictionary IPA
Wiktionary (pt.wiktionary.org)
15,720… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-unified-pronunciation-lexicon.wiktionary-pt-ipa
Portuguese Wiktionary IPA Pronunciations
A comprehensive pronunciation lexicon for Portuguese extracted from pt.wiktionary.org, covering all pages in the category "Entrada com pronúncia (Português)".
Each entry includes IPA transcriptions with regional tags following BCP-47 conventions, plus optional X-SAMPA, part-of-speech, definitions, and etymology.
Dataset Stats
Entries: 17,685
With pronunciations: 15,720 (89%)
With X-SAMPA: 1,563
With POS tags: 14,115 (80%)… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/wiktionary-pt-ipa.dicionario_barranquenho
Dicionário de Barranquenho
Structured lexical dataset derived from the first published dictionary of Barranquenho, a Romance contact language spoken in Barrancos, Portugal. Contains 1,680 entries with Portuguese and Spanish glosses, grammatical categories, semantic fields, source attributions, and synonym cross-references.
Language
Barranquenho (glottocode: barr1245; no ISO 639-3 code assigned at time of publication) is a contact language spoken in the municipality of… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/dicionario_barranquenho.infopedia-pt-ipa
European Portuguese IPA Lexicon — Infopédia
A lightweight word → IPA pronunciation lexicon for European Portuguese,
extracted from Infopédia (Porto Editora). One row
per headword, intended for grapheme-to-phoneme (G2P) work, pronunciation
modelling, and TTS/ASR lexicon building.
Complete crawl. Derived from a graph crawl of Infopédia that ran to
convergence (frontier → 0), covering the dictionary's reachable component.
Contents
Field
Count
Entries… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/infopedia-pt-ipa.portuguese_phonetic_lexicon
📚 Portuguese Phonetic Lexicon Dataset
This dataset contains phonetic and morphological information for Portuguese words, collected from the Portal da Língua Portuguesa. It was generated by scraping the site across multiple Portuguese-speaking regions and dialects.
🌍 Regional Coverage
The dataset includes words as spoken in ten regional variants:
🇵🇹 Lisbon (Standard and Non-Standard)
🇦🇴 Luanda
🇧🇷 Rio de Janeiro (Standard and Non-Standard)
🇧🇷 São Paulo (Standard… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese_phonetic_lexicon.arabic-mantoq-synthetic-g2pportuguese-sentences-synthetic-g2p
Dataset Card for 'TigreGotico/portuguese_g2p'
Dataset Description
Dataset Summary
TigreGotico/portuguese_g2p is a Grapheme-to-Phoneme (G2P) dataset for Portuguese, offering phonetic transcriptions for sentences across ten different regional variants.
It is derived from the portuguese_phonetic_lexicon and is designed to aid in the development of robust Speech Recognition (ASR) and Text-to-Speech (TTS) models that account for dialectal variation in Portuguese.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-sentences-synthetic-g2p.galician_g2p
