SEACrowd/lti_langid_corpus
The LTI LangID corpus is a dataset for language identification. The most recent version, v5, contains training data for 1266 languages, and some (possibly very tiny) amount of text for a total of 1706 languages. This dataloader can only be executed in a BASH environment at the moment. (See https://github.com/SEACrowd/seacrowd-datahub/pull/405)
045
Upload README.md with huggingface_hub
Upload __init__.py with huggingface_hub
Upload lti_langid_corpus.py with huggingface_hub
Upload README.md with huggingface_hub
Upload LICENSE with huggingface_hub
Upload requirements.txt with huggingface_hub
initial commit
