datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mangahausa_aug_lex
title: Lexicon Dataset for the Hausa Language
Dataset with English translation
license: cc-by-nd-4.0
manga-querymultilingual-indic-profane
Dataset Summary
This dataset contains 6081 text entries labeled for safety classification (safe or not safe). The text is multilingual, including native scripts for Malayalam, Hindi, Tamil, and Kannada, as well as their transliterated (romanized) versions. The content ranges from neutral, everyday phrases to highly offensive and profane language. It is suitable for training and evaluating models for tasks like hate speech detection, toxic content filtering, and general text safety… See the full description on the dataset page: https://huggingface.co/datasets/mangalathkedar/multilingual-indic-profane.manga109_emotionSE_Dataset_qnaqna-with-seQNA-chat_appSE-DatasetQuoted_Dataset
