bert_updated
multilingual-pl-bert-is-updated
Introduction
This dataset, derived from the Icelandic Gigaword Corpus, is designed as a more comprehensive alternative to the existing dataset found at
https://huggingface.co/datasets/styletts2-community/multilingual-pl-bert/tree/main/is.
The original dataset, derived from just 52MB of raw text from the Icelandic Wikipedia, was processed using the espeak-ng backend for
normalization and phonemization. However, the Icelandic module of espeak-ng, which has not been updated for over a… See the full description on the dataset page: https://huggingface.co/datasets/grammatek/multilingual-pl-bert-is-updated.multilingual-pl-bert-be-updatedThis is a replacement for the folder https://huggingface.co/datasets/styletts2-community/multilingual-pl-bert/tree/main/be. I haven't changed input_ids, only phonemes. ~95% of all transcriptions were amended where appropriate; the remaining ~5% were copied (with minor changes in notation) from the original version of the dataset.
A few caveats:
My amended transcriptions are based on the outputs of two existing Belarusian G2P tools, corpus.by and bnkorpus.info. They use slightly different IPA… See the full description on the dataset page: https://huggingface.co/datasets/somerandomguyontheweb/multilingual-pl-bert-be-updated.
