datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mecrab-jawiki-word2vec
MeCrab Japanese Word2Vec Vectors
High-quality Japanese word embeddings trained on Wikipedia using MeCrab morphological analyzer.
📊 Dataset Summary
This dataset contains pre-trained Japanese word embeddings optimized for use with MeCrab, a high-performance morphological analyzer.
Key Features:
✅ Trained on Japanese Wikipedia
✅ Zero-copy binary format (MCV1) for fast loading
✅ Compatible with MeCrab Python API
✅ 300-dimensional vectors
✅ ~100,000 vocabulary size… See the full description on the dataset page: https://huggingface.co/datasets/KitaSan/mecrab-jawiki-word2vec.mueller-word2vec-data
