word2vec
word2vec-skipgram-fa-wikipediaword2vec-cbow-fa-wikipediaword2vec-quantizedword2vecdense_encoder-msmarco-bert-base-word2vec256k_emb_updateddense_encoder-msmarco-distilbert-word2vec256k-MLM_445k_emb_updateddense_encoder-msmarco-distilbert-word2vec256k-MLM_210k_emb_updateddense_encoder-msmarco-distilbert-word2vec256k_emb_updated
Datasets
All datasets matching “word2vec”word2vec_arxiveword2vec-google-news-negative-300glove.6B.100d.word2vec.txtword2vec_analogyAdapted from https://github.com/nicholas-leonard/word2vec
mecrab-jawiki-word2vec
MeCrab Japanese Word2Vec Vectors
High-quality Japanese word embeddings trained on Wikipedia using MeCrab morphological analyzer.
📊 Dataset Summary
This dataset contains pre-trained Japanese word embeddings optimized for use with MeCrab, a high-performance morphological analyzer.
Key Features:
✅ Trained on Japanese Wikipedia
✅ Zero-copy binary format (MCV1) for fast loading
✅ Compatible with MeCrab Python API
✅ 300-dimensional vectors
✅ ~100,000 vocabulary size… See the full description on the dataset page: https://huggingface.co/datasets/KitaSan/mecrab-jawiki-word2vec.CHILDES_word2vecAll uploaded word embeddings are trained with Word2Vec on CHILDES data.
available languages:
nor: Norwegian
en: English
