datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
short-text-multi-labeled-emotion-classificationazerbaijani_tweet_emotion_classificationThis dataset contains 150K (train + test) cleaned tweets in Azerbaijani. Tweets were collected in 2021, and filtered and cleaned by following these steps:
Initial data were collected by using twint library. The tool is currently deprecated, cannot be used with new Twitter.
On top of the already filtered data, I applied an additional filter to select Azerbaijani tweets with using fastText language identification model.
Tweets were classified into 3 emotion categories: {positive: 1, negative:… See the full description on the dataset page: https://huggingface.co/datasets/hajili/azerbaijani_tweet_emotion_classification.Indonesian-Emotion-Classificationstory_emotion_classificationemotion_classification
