frequency-list
multilingual-frequency-lists
Multilingual Frequency Lists
This dataset contains multiple word-frequency lists in various languages such as French, Japanese, Spanish, Italian and Portuguese.
Specifically, these are frequency lists of lemmas, meaning, for example, that words like 'run', 'runs' and 'running' are counted together as occurences of the same lemma 'run'.
These frequency lists were generated from ~1GB of subtitles scraped from a variety of Netflix shows and films and parsed using relevant spacy models… See the full description on the dataset page: https://huggingface.co/datasets/joshdavham/multilingual-frequency-lists.english-word-frequency-listBulgarian-Frequency-list
