CoolFace
Datasetpublic

mustafaalkanxgmail/Multilingual-Core-Vocabulary

Multilingual Core Vocabulary 🌍 This dataset contains millions of frequency-sorted, highly accurate words across 19 languages. It is designed to be the ultimate resource for building cross-lingual applications, AI similarity agents, and translation models. Dataset Structure This repository contains two variations of the dataset: Massive_Dataset: Contains over 6 million words. The words were extracted and frequency-sorted from FastText, cleaned from internet noise… See the full description on the dataset page: https://huggingface.co/datasets/mustafaalkanxgmail/Multilingual-Core-Vocabulary.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes78downloads
settings

This repository belongs to mustafaalkanxgmail on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameMultilingual-Core-Vocabulary
visibilitypublic
licencecc-by-4.0
gatedno
ownermustafaalkanxgmail
Account settings
mustafaalkanxgmail/Multilingual-Core-Vocabulary · CoolFace