CoolFace
14 results

PL-BERT

styletts2-community /multilingual-pl-bertAttribution: Wikipedia.org text100K<n<1M19 likes4.2k downloads3y agoHugging Faceoddadmix /arabic-pl-bert Dataset Card for "arabic-pl-bert" More Information needed text1M<n<10M0 likes82 downloads5mo agoHugging Facehon9kon9ize /yue-wiki-pl-bert Yue-Wiki-PL-BERT Dataset Overview This dataset contains processed text data from Cantonese Wikipedia articles, specifically formatted for training or fine-tuning BERT-like models for Cantonese language processing. The dataset is created by hon9kon9ize and contains approximately 176,177 rows of training data. Description The Yue-Wiki-PL-BERT dataset is a structured collection of Cantonese text data extracted from Wikipedia, with each entry containing: id: A… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue-wiki-pl-bert.text100K<n<1M1 likes59 downloads1y agoHugging Facegrammatek /multilingual-pl-bert-is-updated Introduction This dataset, derived from the Icelandic Gigaword Corpus, is designed as a more comprehensive alternative to the existing dataset found at https://huggingface.co/datasets/styletts2-community/multilingual-pl-bert/tree/main/is. The original dataset, derived from just 52MB of raw text from the Icelandic Wikipedia, was processed using the espeak-ng backend for normalization and phonemization. However, the Icelandic module of espeak-ng, which has not been updated for over a… See the full description on the dataset page: https://huggingface.co/datasets/grammatek/multilingual-pl-bert-is-updated.text100K<n<1M0 likes57 downloads3y agoHugging Facemesolitica /PL-BERT-MS PL-BERT-MS dataset Combine wikimedia/wikipedia on /20231101.ms with news dataset. Tokenizer from mesolitica/PL-BERT-MS. Source code All source code at https://github.com/mesolitica/PL-BERT-MS text1M<n<10M0 likes26 downloads2y agoHugging Facesomerandomguyontheweb /multilingual-pl-bert-be-updatedThis is a replacement for the folder https://huggingface.co/datasets/styletts2-community/multilingual-pl-bert/tree/main/be. I haven't changed input_ids, only phonemes. ~95% of all transcriptions were amended where appropriate; the remaining ~5% were copied (with minor changes in notation) from the original version of the dataset. A few caveats: My amended transcriptions are based on the outputs of two existing Belarusian G2P tools, corpus.by and bnkorpus.info. They use slightly different IPA… See the full description on the dataset page: https://huggingface.co/datasets/somerandomguyontheweb/multilingual-pl-bert-be-updated.text100K<n<1M0 likes18 downloads3y agoHugging Face