sentencepiece
wmt21-train-tokenized-sentencepiecesentencepiece-wikitext-103Caution ⚠️ : This dataset is already tokenized using the sentencepiece tokenizer from "ibm-aimc/sigma-moe-small".
This dataset is an adapted version from https://huggingface.co/datasets/EleutherAI/wikitext_document_level which in turn is an adapted version of the dataset in https://huggingface.co/datasets/wikitext .
Human_DNA_v0_SentencepieceTokenized_vocab30kbabylm-sentencepiece16k-512-encodedHuman_DNA_v0_SentencepieceTokenized_vocab10kCopus_for_SentencePiece
#2024/5/20#create @ K_Shioiri
SentencePiece用の学習Copus
http://www.lsta.media.kyoto-u.ac.jp/resource/data/wikitext-ja/home.htmlCopus_wiki_good:日本語wikipedia 優秀な記事Copus_wiki_featured:日本語wikipedia 良好な記事
https://data.statmt.org/cc-100/
Copus_cc100_ja:日本語CC100original:458,387,942 data・[]{}<>【】に囲まれた語彙の削除・アドレス(@)を含む文、httmを含む文の削除・10words未満、200words以上の文の削除-> transform: 383,904,390 dataCopus_Coarse_cc100_ja:日本語CC100上記の間引き(about 1/10) -> 38,390,439 data
using… See the full description on the dataset page: https://huggingface.co/datasets/SaltyCedar/Copus_for_SentencePiece.
