CoolFace
Datasetpublicgated

junwatu/javanese-multilingual-lexicon

Javanese Multilingual Lexicon A multilingual lexicon dataset of 202,912 entries derived from Sastra.org, a digital archive of Javanese literary heritage maintained by Yayasan Sastra Lestari. The dataset compiles 32 historical lexicographic sources spanning from the 1830s to 2010, covering Javanese, Old Javanese (Kawi), Indonesian, Dutch, English, and French. Dataset Structure The dataset is provided as JSONL files. The full corpus is in data/leksikon_all.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/junwatu/javanese-multilingual-lexicon.

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
1likes17downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.