Multilingual BERT
bert-base-multilingual-uncasedbert-base-multilingual-casedbert-base-multilingual-uncased-sentimentllmlingua-2-bert-base-multilingual-cased-meetingbankbert-base-multilingual-cased-ner-hrlbert-base-multilingual-uncased-sentimentbert-base-multilingual-cased-ner-hrlbert-base-multilingual-cased-language-detection
multilingual-pl-bertAttribution: Wikipedia.org
multilingual-pl-bert-is-updated
Introduction
This dataset, derived from the Icelandic Gigaword Corpus, is designed as a more comprehensive alternative to the existing dataset found at
https://huggingface.co/datasets/styletts2-community/multilingual-pl-bert/tree/main/is.
The original dataset, derived from just 52MB of raw text from the Icelandic Wikipedia, was processed using the espeak-ng backend for
normalization and phonemization. However, the Icelandic module of espeak-ng, which has not been updated for over a… See the full description on the dataset page: https://huggingface.co/datasets/grammatek/multilingual-pl-bert-is-updated.bert-base-multilingual-cased-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
multilingual-pl-bert-be-updatedThis is a replacement for the folder https://huggingface.co/datasets/styletts2-community/multilingual-pl-bert/tree/main/be. I haven't changed input_ids, only phonemes. ~95% of all transcriptions were amended where appropriate; the remaining ~5% were copied (with minor changes in notation) from the original version of the dataset.
A few caveats:
My amended transcriptions are based on the outputs of two existing Belarusian G2P tools, corpus.by and bnkorpus.info. They use slightly different IPA… See the full description on the dataset page: https://huggingface.co/datasets/somerandomguyontheweb/multilingual-pl-bert-be-updated.multilingual-bert-toc-95k-dataset
Dataset Details
Dataset Description
Contains line-by-line sequences from human-annotated legal/government documents and their corresponding labels.
Line-by-line examples derived from DocLayNet dataset
Dataset Creation
Notebook displaying how dataset was created can be accessed here
bert_base_multilingual_cased_predictions
