datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stop_wordsnagisa_stopwords
Japanese Stopwords for nagisa
This dataset is the Japanese stopwords list built into nagisa (v0.2.12+). It is published here on Hugging Face for easy access and reproducibility.
Overview
Language
Japanese
Size
147 words
Source
CC-100, Wikipedia
License
MIT
Dataset Description
This dataset contains 147 frequently used Japanese words extracted from large-scale corpora. Each word is annotated with its part-of-speech (POS) tag according… See the full description on the dataset page: https://huggingface.co/datasets/taishi-i/nagisa_stopwords.stopwordsstopwordsindo-stopwordsmodel_test_stopwords
