CoolFace
20 results

Multilingual BERT

styletts2-community /multilingual-pl-bertAttribution: Wikipedia.org text100K<n<1M19 likes4.2k downloads3y agoHugging Facegrammatek /multilingual-pl-bert-is-updated Introduction This dataset, derived from the Icelandic Gigaword Corpus, is designed as a more comprehensive alternative to the existing dataset found at https://huggingface.co/datasets/styletts2-community/multilingual-pl-bert/tree/main/is. The original dataset, derived from just 52MB of raw text from the Icelandic Wikipedia, was processed using the espeak-ng backend for normalization and phonemization. However, the Icelandic module of espeak-ng, which has not been updated for over a… See the full description on the dataset page: https://huggingface.co/datasets/grammatek/multilingual-pl-bert-is-updated.text100K<n<1M0 likes59 downloads3y agoHugging Facetoksuitebackup /bert-base-multilingual-cased-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes58 downloads10mo agoHugging Facesomerandomguyontheweb /multilingual-pl-bert-be-updatedThis is a replacement for the folder https://huggingface.co/datasets/styletts2-community/multilingual-pl-bert/tree/main/be. I haven't changed input_ids, only phonemes. ~95% of all transcriptions were amended where appropriate; the remaining ~5% were copied (with minor changes in notation) from the original version of the dataset. A few caveats: My amended transcriptions are based on the outputs of two existing Belarusian G2P tools, corpus.by and bnkorpus.info. They use slightly different IPA… See the full description on the dataset page: https://huggingface.co/datasets/somerandomguyontheweb/multilingual-pl-bert-be-updated.text100K<n<1M0 likes18 downloads3y agoHugging Facepruhtopia /multilingual-bert-toc-95k-dataset Dataset Details Dataset Description Contains line-by-line sequences from human-annotated legal/government documents and their corresponding labels. Line-by-line examples derived from DocLayNet dataset Dataset Creation Notebook displaying how dataset was created can be accessed here texttext-classification10K<n<100K0 likes12 downloads2y agoHugging Faceartianand /bert_base_multilingual_cased_predictionstabular10K<n<100K0 likes7 downloads10mo agoHugging Face