CoolFace
20 results

langid

strombergnlp /nordic_langidAutomatic language identification is a challenging problem. Discriminating between closely related languages is especially difficult. This paper presents a machine learning approach for automatic language identification for the Nordic languages, which often suffer miscategorisation by existing state-of-the-art tools. Concretely we will focus on discrimination between six Nordic languages: Danish, Swedish, Norwegian (Nynorsk), Norwegian (Bokmål), Faroese and Icelandic. This is the data for the tasks. Two variants are provided: 10K and 50K, with holding 10,000 and 50,000 examples for each language respectively.texttext-classification100K<n<1M5 likes344 downloads4y agoHugging FaceaakashMeghwar01 /sindhi-corpus-langid-cleantext100K<n<1M0 likes55 downloads3mo agoHugging FaceSEACrowd /lti_langid_corpusThe LTI LangID corpus is a dataset for language identification. The most recent version, v5, contains training data for 1266 languages, and some (possibly very tiny) amount of text for a total of 1706 languages. This dataloader can only be executed in a BASH environment at the moment. (See https://github.com/SEACrowd/seacrowd-datahub/pull/405)0 likes45 downloads2y agoHugging FaceMaleeshaK /Sinhala-Script-LangID-Benchmarktext10K<n<100K0 likes33 downloads19d agoHugging Facerachel2999 /lang_identtextn<1K0 likes30 downloads3y agoHugging Facekardosdrur /scandi-langid Dataset Card for "scandi-langid" More Information needed text100K<n<1M0 likes27 downloads3y agoHugging Face