CoolFace
Datasetpublicgated

junwatu/javanese-multilingual-lexicon

Javanese Multilingual Lexicon A multilingual lexicon dataset of 202,912 entries derived from Sastra.org, a digital archive of Javanese literary heritage maintained by Yayasan Sastra Lestari. The dataset compiles 32 historical lexicographic sources spanning from the 1830s to 2010, covering Javanese, Old Javanese (Kawi), Indonesian, Dutch, English, and French. Dataset Structure The dataset is provided as JSONL files. The full corpus is in data/leksikon_all.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/junwatu/javanese-multilingual-lexicon.

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
1likes16downloads
settings

This repository belongs to junwatu on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namejavanese-multilingual-lexicon
visibilitypublic
licencecc-by-nc-4.0
gatedyes
ownerjunwatu
Account settings
junwatu/javanese-multilingual-lexicon · CoolFace