CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01papluca /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.texttext-classification10K<n<100K70 likes4.2k downloads4y agoHugging Face02unklefedor /language-identificationtext100K<n<1M1 likes154 downloads3y agoHugging Face03pranavagrawal /Language-Identificationtext1M<n<10M0 likes50 downloads2y agoHugging Face04mlexplorer008 /south_african_language_identificationtext10K<n<100K1 likes30 downloads2y agoHugging Face05chiragkolte01 /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/chiragkolte01/language-identification.texttext-classification10K<n<100K0 likes22 downloads5mo agoHugging Face06borrore /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/borrore/language-identification.texttext-classification10K<n<100K0 likes18 downloads6mo agoHugging Face07sameeramin /code-switched-language-identificationtext100K<n<1M0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.