CoolFace
Datasetpublic

SEACrowd/id_multilabel_hs

The ID_MULTILABEL_HS dataset is collection of 13,169 tweets in Indonesian language, designed for hate speech detection NLP task. This dataset is combination from previous research and newly crawled data from Twitter. This is a multilabel dataset with label details as follows: -HS : hate speech label; -Abusive : abusive language label; -HS_Individual : hate speech targeted to an individual; -HS_Group : hate speech targeted to a group; -HS_Religion : hate speech related to religion/creed; -HS_Race : hate speech related to race/ethnicity; -HS_Physical : hate speech related to physical/disability; -HS_Gender : hate speech related to gender/sexual orientation; -HS_Gender : hate related to other invective/slander; -HS_Weak : weak hate speech; -HS_Moderate : moderate hate speech; -HS_Strong : strong hate speech.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes159downloads
README.md95 linesDownload Raw Back to root
1 2---3language: 4- ind5pretty_name: Id Multilabel Hs6task_categories: 7- aspect-based-sentiment-analysis8tags: 9- aspect-based-sentiment-analysis10---11 12The ID_MULTILABEL_HS dataset is collection of 13,169 tweets in Indonesian language,13designed for hate speech detection NLP task. This dataset is combination from previous research and newly crawled data from Twitter.14This is a multilabel dataset with label details as follows:15-HS : hate speech label;16-Abusive : abusive language label;17-HS_Individual : hate speech targeted to an individual;18-HS_Group : hate speech targeted to a group;19-HS_Religion : hate speech related to religion/creed;20-HS_Race : hate speech related to race/ethnicity;21-HS_Physical : hate speech related to physical/disability;22-HS_Gender : hate speech related to gender/sexual orientation;23-HS_Gender : hate related to other invective/slander;24-HS_Weak : weak hate speech;25-HS_Moderate : moderate hate speech;26-HS_Strong : strong hate speech.27 28 29## Languages30 31ind32 33## Supported Tasks34 35Aspect Based Sentiment Analysis36 37## Dataset Usage38### Using `datasets` library39```40from datasets import load_dataset41dset = datasets.load_dataset("SEACrowd/id_multilabel_hs", trust_remote_code=True)42```43### Using `seacrowd` library44```import seacrowd as sc45# Load the dataset using the default config46dset = sc.load_dataset("id_multilabel_hs", schema="seacrowd")47# Check all available subsets (config names) of the dataset48print(sc.available_config_names("id_multilabel_hs"))49# Load the dataset using a specific config50dset = sc.load_dataset_by_config_name(config_name="<config_name>")51```52 53More details on how to load the `seacrowd` library can be found [here](https://github.com/SEACrowd/seacrowd-datahub?tab=readme-ov-file#how-to-use).54 55 56## Dataset Homepage57 58[https://aclanthology.org/W19-3506/](https://aclanthology.org/W19-3506/)59 60## Dataset Version61 62Source: 1.0.0. SEACrowd: 2024.06.20.63 64## Dataset License65 66Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International67 68## Citation69 70If you are using the **Id Multilabel Hs** dataloader in your work, please cite the following:71```72@inproceedings{ibrohim-budi-2019-multi,73    title = "Multi-label Hate Speech and Abusive Language Detection in {I}ndonesian {T}witter",74    author = "Ibrohim, Muhammad Okky  and75      Budi, Indra",76    booktitle = "Proceedings of the Third Workshop on Abusive Language Online",77    month = aug,78    year = "2019",79    address = "Florence, Italy",80    publisher = "Association for Computational Linguistics",81    url = "https://aclanthology.org/W19-3506",82    doi = "10.18653/v1/W19-3506",83    pages = "46--57",84}85 86 87@article{lovenia2024seacrowd,88    title={SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages}, 89    author={Holy Lovenia and Rahmad Mahendra and Salsabil Maulana Akbar and Lester James V. Miranda and Jennifer Santoso and Elyanah Aco and Akhdan Fadhilah and Jonibek Mansurov and Joseph Marvin Imperial and Onno P. Kampman and Joel Ruben Antony Moniz and Muhammad Ravi Shulthan Habibi and Frederikus Hudi and Railey Montalan and Ryan Ignatius and Joanito Agili Lopo and William Nixon and Börje F. Karlsson and James Jaya and Ryandito Diandaru and Yuze Gao and Patrick Amadeus and Bin Wang and Jan Christian Blaise Cruz and Chenxi Whitehouse and Ivan Halim Parmonangan and Maria Khelli and Wenyu Zhang and Lucky Susanto and Reynard Adha Ryanda and Sonny Lazuardi Hermawan and Dan John Velasco and Muhammad Dehan Al Kautsar and Willy Fitra Hendria and Yasmin Moslem and Noah Flynn and Muhammad Farid Adilazuarda and Haochen Li and Johanes Lee and R. Damanhuri and Shuo Sun and Muhammad Reza Qorib and Amirbek Djanibekov and Wei Qi Leong and Quyet V. Do and Niklas Muennighoff and Tanrada Pansuwan and Ilham Firdausi Putra and Yan Xu and Ngee Chia Tai and Ayu Purwarianti and Sebastian Ruder and William Tjhi and Peerat Limkonchotiwat and Alham Fikri Aji and Sedrick Keh and Genta Indra Winata and Ruochen Zhang and Fajri Koto and Zheng-Xin Yong and Samuel Cahyawijaya},90    year={2024},91    eprint={2406.10118},92    journal={arXiv preprint arXiv: 2406.10118}93}94 95```