datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
word2vec_arxiveglove.6B.100d.word2vec.txtword2vec_analogyAdapted from https://github.com/nicholas-leonard/word2vec
mecrab-jawiki-word2vec
MeCrab Japanese Word2Vec Vectors
High-quality Japanese word embeddings trained on Wikipedia using MeCrab morphological analyzer.
📊 Dataset Summary
This dataset contains pre-trained Japanese word embeddings optimized for use with MeCrab, a high-performance morphological analyzer.
Key Features:
✅ Trained on Japanese Wikipedia
✅ Zero-copy binary format (MCV1) for fast loading
✅ Compatible with MeCrab Python API
✅ 300-dimensional vectors
✅ ~100,000 vocabulary size… See the full description on the dataset page: https://huggingface.co/datasets/KitaSan/mecrab-jawiki-word2vec.squad_v2_word2vec_augword2vecmueller-word2vec-datapruned-word2vecWord2Vec_400dim
