CoolFace
Datasetpublic

KeisukeMiyamoto/lambda-corpus

lambda-corpus lambda-corpus is a Japanese text corpus for language model pretraining. It combines openly available Japanese datasets into a consistent format and provides predefined train, validation, and test splits. Purpose The dataset is intended for pretraining of Japanese language models. It contains web documents, Wikipedia-derived text, academic grant records, and synthetic question-answer text. Source Data Source Rows Tokens License… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/lambda-corpus.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes276downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

KeisukeMiyamoto/lambda-corpus · main · files are served by the source, never re-hosted here