datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
profanityThis dataset is originaly from https://github.com/vzhou842/profanity-check
multilingual-swear-profanity
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/miklgr500/jigsaw-multilingual-swear-profanity
Multilingual swear profanity
Current dataset consist of swear profanity on six languages:
French (fr)Turkish (tr)Italian (it)Russian (ru)Spanish (es)Portugalian (pt)
Sources:
Italian Swear Words, Phrases, Curses, Insults, Slang, Colloquialisms and Expletives!Italian swear wordsItalian profanity (wiki)Turkish/SlangTurkish Slang DictionaryTurkish Swear… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-swear-profanity.profanity
