laurievb/open-lid-dataset
Dataset Card for "open-lid-dataset" Dataset Summary The OpenLID dataset covers 201 languages and is designed for training language identification models. The majority of the source datasets were derived from news sites, Wikipedia, or religious text, though some come from other domains (e.g. transcribed conversations, literature, or social media). A sample of each language in each source was manually audited to check it was in the attested language (see the paper)… See the full description on the dataset page: https://huggingface.co/datasets/laurievb/open-lid-dataset.
52.3k
