CoolFace
Datasetpublic

joshdavham/multilingual-frequency-lists

Multilingual Frequency Lists This dataset contains multiple word-frequency lists in various languages such as French, Japanese, Spanish, Italian and Portuguese. Specifically, these are frequency lists of lemmas, meaning, for example, that words like 'run', 'runs' and 'running' are counted together as occurences of the same lemma 'run'. These frequency lists were generated from ~1GB of subtitles scraped from a variety of Netflix shows and films and parsed using relevant spacy… See the full description on the dataset page: https://huggingface.co/datasets/joshdavham/multilingual-frequency-lists.

sourceHugging Faceagpl-3.0updated 5mo agoView on Hugging Face
1likes78downloads

joshdavham/multilingual-frequency-lists · main · files are served by the source, never re-hosted here