PL-BERT
Datasets
All datasets matching “PL-BERT”multilingual-pl-bertAttribution: Wikipedia.org
arabic-pl-bert
Dataset Card for "arabic-pl-bert"
More Information needed
yue-wiki-pl-bert
Yue-Wiki-PL-BERT Dataset
Overview
This dataset contains processed text data from Cantonese Wikipedia articles, specifically formatted for training or fine-tuning BERT-like models for Cantonese language processing. The dataset is created by hon9kon9ize and contains approximately 176,177 rows of training data.
Description
The Yue-Wiki-PL-BERT dataset is a structured collection of Cantonese text data extracted from Wikipedia, with each entry containing:
id: A… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue-wiki-pl-bert.multilingual-pl-bert-is-updated
Introduction
This dataset, derived from the Icelandic Gigaword Corpus, is designed as a more comprehensive alternative to the existing dataset found at
https://huggingface.co/datasets/styletts2-community/multilingual-pl-bert/tree/main/is.
The original dataset, derived from just 52MB of raw text from the Icelandic Wikipedia, was processed using the espeak-ng backend for
normalization and phonemization. However, the Icelandic module of espeak-ng, which has not been updated for over a… See the full description on the dataset page: https://huggingface.co/datasets/grammatek/multilingual-pl-bert-is-updated.PL-BERT-MS
PL-BERT-MS dataset
Combine wikimedia/wikipedia on /20231101.ms with news dataset.
Tokenizer from mesolitica/PL-BERT-MS.
Source code
All source code at https://github.com/mesolitica/PL-BERT-MS
multilingual-pl-bert-be-updatedThis is a replacement for the folder https://huggingface.co/datasets/styletts2-community/multilingual-pl-bert/tree/main/be. I haven't changed input_ids, only phonemes. ~95% of all transcriptions were amended where appropriate; the remaining ~5% were copied (with minor changes in notation) from the original version of the dataset.
A few caveats:
My amended transcriptions are based on the outputs of two existing Belarusian G2P tools, corpus.by and bnkorpus.info. They use slightly different IPA… See the full description on the dataset page: https://huggingface.co/datasets/somerandomguyontheweb/multilingual-pl-bert-be-updated.
