datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
portuguese_phonetic_lexicon
📚 Portuguese Phonetic Lexicon Dataset
This dataset contains phonetic and morphological information for Portuguese words, collected from the Portal da Língua Portuguesa. It was generated by scraping the site across multiple Portuguese-speaking regions and dialects.
🌍 Regional Coverage
The dataset includes words as spoken in ten regional variants:
🇵🇹 Lisbon (Standard and Non-Standard)
🇦🇴 Luanda
🇧🇷 Rio de Janeiro (Standard and Non-Standard)
🇧🇷 São Paulo (Standard… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese_phonetic_lexicon.chew_lexical
Dataset Card for Dataset Name
This is the lexical/no-overlapping split of the CHEW dataset(CHEW: A Dataset of CHanging Events in Wikipedia).
Dataset Details
Dataset Description
This dataset is the Lexical/No-overlapping split of the CHEW Dataset,where CHEW stands for CHanging Events in Wikipedia. It contains Wikipedia titles, text in two timestamped versions and Binary Label showing Change(1) or No change(0). Change here means there has been informationm… See the full description on the dataset page: https://huggingface.co/datasets/hsuvaskakoty/chew_lexical.
