ce-lery/merged-corpus
Merged Corpus Welcome to this repository.This dataset is japanese corpus that includes wiki, wikibooks, wikiversity, cc100, and oscar2109. Getting Started If you want to use this, please run as follows.This process takes about 3 hours. mkdir -p pretrain/input/ cd pretrain/input/ GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/ce-lery/merged-corpus.git cd merged-corpus git lfs pull bash merge_train.sh
031
