CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01candradhipa /Language-Detectiontext10K<n<100K0 likes1.6k downloads2y agoHugging Face02sakthivinash /Language_Detection Language_Detection - Multilingual Text Classification Dataset This dataset is a collection of multilingual text samples designed for training and predicting languages in Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), and Data Science (DS) applications. It contains labeled data that associates text samples with their respective languages, enabling language detection and classification tasks. Dataset Overview The dataset consists of two columns:… See the full description on the dataset page: https://huggingface.co/datasets/sakthivinash/Language_Detection.text10K<n<100K0 likes1.5k downloads2y agoHugging Face03simoneteglia /europarl_for_language_detection_10ktext100K<n<1M0 likes739 downloads3y agoHugging Face04FrancophonIA /language_detection [!NOTE] Dataset origin: https://www.kaggle.com/datasets/basilb2s/language-detection It's a small language detection dataset. This dataset consists of text details for 17 different languages, ie, you will be able to create an NLP model for predicting 17 different language.. text10K<n<100K0 likes337 downloads1y agoHugging Face05MoazIrfan /language-detectiontext10K<n<100K1 likes322 downloads9mo agoHugging Face06Toygar /turkish-offensive-language-detection Dataset Summary This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.tabulartext-classification10K<n<100K20 likes171 downloads3y agoHugging Face07dirtycomputer /Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_Languagetabular10K<n<100K0 likes105 downloads3y agoHugging Face08md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes39 downloads3y agoHugging Face09kumarsujit474 /augmented-singlish-language-detection-corpustext10K<n<100K0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.