CoolFace
Datasetpublic

laurievb/open-lid-dataset

Dataset Card for "open-lid-dataset" Dataset Summary The OpenLID dataset covers 201 languages and is designed for training language identification models. The majority of the source datasets were derived from news sites, Wikipedia, or religious text, though some come from other domains (e.g. transcribed conversations, literature, or social media). A sample of each language in each source was manually audited to check it was in the attested language (see the paper)… See the full description on the dataset page: https://huggingface.co/datasets/laurievb/open-lid-dataset.

sourceHugging Faceotherupdated 3y agoView on Hugging Face
5likes2.3kdownloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
laurievb/open-lid-dataset · CoolFace