CoolFace
15 results

sentencepiece

Geonwoohong /wmt21-train-tokenized-sentencepiece1M<n<10M0 likes91 downloads11mo agoHugging Faceibm-aimc /sentencepiece-wikitext-103Caution ⚠️ : This dataset is already tokenized using the sentencepiece tokenizer from "ibm-aimc/sigma-moe-small". This dataset is an adapted version from https://huggingface.co/datasets/EleutherAI/wikitext_document_level which in turn is an adapted version of the dataset in https://huggingface.co/datasets/wikitext . 100K<n<1M0 likes49 downloads3y agoHugging FaceVlasta /Human_DNA_v0_SentencepieceTokenized_vocab30k100K<n<1M0 likes33 downloads4y agoHugging Facemiguelcsx /babylm-sentencepiece16k-512-encoded0 likes29 downloads3mo agoHugging FaceVlasta /Human_DNA_v0_SentencepieceTokenized_vocab10k100K<n<1M0 likes25 downloads4y agoHugging FaceSaltyCedar /Copus_for_SentencePiece #2024/5/20#create @ K_Shioiri SentencePiece用の学習Copus http://www.lsta.media.kyoto-u.ac.jp/resource/data/wikitext-ja/home.htmlCopus_wiki_good:日本語wikipedia 優秀な記事Copus_wiki_featured:日本語wikipedia 良好な記事 https://data.statmt.org/cc-100/ Copus_cc100_ja:日本語CC100original:458,387,942 data・[]{}<>【】に囲まれた語彙の削除・アドレス(@)を含む文、httmを含む文の削除・10words未満、200words以上の文の削除-> transform: 383,904,390 dataCopus_Coarse_cc100_ja:日本語CC100上記の間引き(about 1/10) -> 38,390,439 data using… See the full description on the dataset page: https://huggingface.co/datasets/SaltyCedar/Copus_for_SentencePiece.text10M<n<100M0 likes3 downloads2y agoHugging Face