CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hatakeyama-llm-team /PMC Data collected from PMC Only CC-BY, CC-BY-SA licenses are included. For all records, check the jsonl files in the data folder text100K<n<1M2 likes13k downloads2y agoHugging Face02nayohan /korean-hate-speechreference: https://github.com/kocohub/korean-hate-speech @inproceedings{moon-etal-2020-beep, title = "{BEEP}! {K}orean Corpus of Online News Comments for Toxic Speech Detection", author = "Moon, Jihyung and Cho, Won Ik and Lee, Junbum", booktitle = "Proceedings of the Eighth International Workshop on Natural Language Processing for Social Media", month = jul, year = "2020", address = "Online", publisher = "Association for Computational Linguistics"… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/korean-hate-speech.text1K<n<10K2 likes4.1k downloads2y agoHugging Face03ucberkeley-dlab /measuring-hate-speech Dataset card for Measuring Hate Speech This is a public release of the dataset described in Kennedy et al. (2020) and Sachdeva et al. (2022), consisting of 39,565 comments annotated by 7,912 annotators, for 135,556 combined rows. The primary outcome variable is the "hate speech score" but the 10 constituent ordinal labels (sentiment, (dis)respect, insult, humiliation, inferior status, violence, dehumanization, genocide, attack/defense, hate speech benchmark) can also be treated as… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/measuring-hate-speech.tabulartext-classification100K<n<1M54 likes2.1k downloads9mo agoHugging Face04tdavidson /hate_speech_offensive Dataset Card for [Dataset Name] Dataset Summary An annotated dataset for hate speech and offensive language detection on tweets. Supported Tasks and Leaderboards [More Information Needed] Languages English (en) Dataset Structure Data Instances { "count": 3, "hate_speech_annotation": 0, "offensive_language_annotation": 0, "neither_annotation": 3, "label": 2, # "neither" "tweet": "!!! RT @mayasolovely: As a woman you… See the full description on the dataset page: https://huggingface.co/datasets/tdavidson/hate_speech_offensive.tabulartext-classification10K<n<100K42 likes2.1k downloads3y agoHugging Face05yatin-superintelligence /White-Hat-Security-Agent-Prompts-600K White Hat Security Agent Prompts 600K Overview The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios. Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.texttext-generation100K<n<1M21 likes1.9k downloads6mo agoHugging Face06cs5242-hateful-memes /hateful-memes-data Hateful Memes (CS5242 submission mirror) Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020) used for reproducibility of our CS5242 (NUS) submission. Contents img/ — 10,000 PNG images of memes train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540), test_seen.jsonl (1,000), test_unseen.jsonl (2,000) Provenance This mirror merges two existing mirrors of the original Meta release: Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/cs5242-hateful-memes/hateful-memes-data.imageimage-classification10K<n<100K2 likes1.2k downloads5mo agoHugging Face07tweets-hate-speech-detection /tweets_hate_speech_detection Dataset Card for Tweets Hate Speech Detection Dataset Summary The objective of this task is to detect hate speech in tweets. For the sake of simplicity, we say a tweet contains hate speech if it has a racist or sexist sentiment associated with it. So, the task is to classify racist or sexist tweets from other tweets. Formally, given a training sample of tweets and labels, where label ‘1’ denotes the tweet is racist/sexist and label ‘0’ denotes the tweet is not… See the full description on the dataset page: https://huggingface.co/datasets/tweets-hate-speech-detection/tweets_hate_speech_detection.texttext-classification10K<n<100K18 likes825 downloads2y agoHugging Face08team-hatakeyama-phase2 /ndlj_tosho_1 国会図書館に収蔵される著作権切れのデータです textn<1K0 likes713 downloads2y agoHugging Face09astronolan /galaxy-mentions-hats Galaxy Mentions HATS A HATS catalog of 43,546 resolved literature mentions from astronolan/galaxy-mentions, prepared for efficient spatial crossmatching with the Multimodal Universe HATS catalogs. Each row is a mention in a paper, not a deduplicated astronomical object. Multiple rows may therefore describe the same galaxy. The optional wiki_entity_id groups mentions using the current 1 arcsecond connected-components build, while mention_id remains the unique row identifier.… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/galaxy-mentions-hats.tabular10K<n<100K1 likes510 downloads2mo agoHugging Face10hatemestinbejaia /miracl-arabictext10M<n<100M0 likes437 downloads2mo agoHugging Face11hatakeyama-llm-team /japanese2010 日本語ウェブコーパス2010 こちらのデータをhuggingfaceにアップロードしたものです。 2009 年度における著作権法の改正(平成21年通常国会 著作権法改正等について | 文化庁)に基づき,情報解析研究への利用に限って利用可能です。 形態素解析を用いて、自動で句点をつけました。 変換コード 変換スクリプト 形態素解析など text1M<n<10M3 likes413 downloads3y agoHugging Face12hatakeyama-llm-team /CommonCrawl_wet_v2text100K<n<1M1 likes387 downloads3y agoHugging Face13community-datasets /roman_urdu_hate_speech Dataset Card for roman_urdu_hate_speech Dataset Summary The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.texttext-classification10K<n<100K3 likes319 downloads2y agoHugging Face14Akajackson /synth_hathor_20000image10K<n<100K0 likes274 downloads2y agoHugging Face15apple /hat HAT: Hallucination Annotation for Translation 🧭 Table of Contents Overview Usage Data Creation Process Data Statistics Dataset Structure Paper Abstract Citation License 📘 Overview HAT (Hallucination Annotation for Translation) is a large-scale dataset for hallucination detection in machine translation (MT).It is released as part of our publication at ACL 2026 (paper). 350,959 span-level annotated samples 38 language pairs ~8,000-10,000… See the full description on the dataset page: https://huggingface.co/datasets/apple/hat.tabulartranslation100K<n<1M19 likes249 downloads3mo agoHugging Face16mteb /MMSoc_HatefulMemesimage10K<n<100K0 likes239 downloads8mo agoHugging Face17rezacsedu /bn_hate_speech Dataset Card for Bengali Hate Speech Dataset Dataset Summary The Bengali Hate Speech Dataset is a Bengali-language dataset of news articles collected from various Bengali media sources and categorized based on the type of hate in the text. The dataset was created to provide greater support for under-resourced languages like Bengali on NLP tasks, and serves as a benchmark for multiple types of classification tasks. Supported Tasks and Leaderboards topic… See the full description on the dataset page: https://huggingface.co/datasets/rezacsedu/bn_hate_speech.texttext-classification1K<n<10K3 likes226 downloads3y agoHugging Face18hatemestinbejaia /STCALIR_Synthetic-Test-Collectiontext100K<n<1M0 likes224 downloads6mo agoHugging Face19dffeewew /hateful_memes Facebook Hateful Memes Dataset Complete version of the Hateful Memes Challenge dataset (Kiela et al., 2020) with all images included. Dataset Description Hateful memes combine individually benign images and text to produce hateful content. The hate lives in the interaction between modalities, making this one of the hardest content moderation benchmarks. The dataset includes confounders: meme pairs that share the same text (or image) but carry opposite labels, forcing… See the full description on the dataset page: https://huggingface.co/datasets/dffeewew/hateful_memes.imageimage-classification10K<n<100K0 likes218 downloads7mo agoHugging Face20piuba-bigdata /contextualized_hate_speech Contextualized Hate Speech: A dataset of comments in news outlets on Twitter Dataset Summary This dataset is a collection of tweets that were posted in response to news articles from five specific Argentinean news outlets: Clarín, Infobae, La Nación, Perfil and Crónica, during the COVID-19 pandemic. The comments were analyzed for hate speech across eight different characteristics: against women, racist content, class hatred, against LGBTQ+ individuals, against physical… See the full description on the dataset page: https://huggingface.co/datasets/piuba-bigdata/contextualized_hate_speech.texttext-classification10K<n<100K8 likes203 downloads3y agoHugging Face21sayvan /vqa_hataw_v1image10K<n<100K0 likes202 downloads7mo agoHugging Face22aiatums /codemixed-id-hate-speech Code-mixed Indonesian Hate Speech Dataset Manually annotated hate speech dataset for Indonesian-Javanese and Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations. texttext-classification10K<n<100K0 likes197 downloads4mo agoHugging Face23Lots-of-LoRAs /task905_hate_speech_offensive_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task905_hate_speech_offensive_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task905_hate_speech_offensive_classification.texttext-generation1K<n<10K0 likes188 downloads2y agoHugging Face24nirmalendu01 /hateXplain_filteredtabular10K<n<100K0 likes183 downloads2y agoHugging Face25community-datasets /hate_speech_pl Dataset Card for HateSpeechPl Dataset Summary The dataset was created to analyze the possibility of automating the recognition of hate speech in Polish. It was collected from the Polish forums and represents various types and degrees of offensive language, expressed towards minorities. The original dataset is provided as an export of MySQL tables, what makes it hard to load. Due to that, it was converted to CSV and put to a Github repository. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hate_speech_pl.tabulartext-classification10K<n<100K4 likes159 downloads2y agoHugging Face26suwaimyo /hatespeech-ind-multilabelclassification HateSpeech_ind_MultiLabelClassification Deduplicated copy of kornwtp/hatespeech-ind-multilabelclassification. Splits split rows train 13,014 text10K<n<100K0 likes136 downloads24d agoHugging Face27mteb /HateSpeechPortugueseClassification HateSpeechPortugueseClassification An MTEB dataset Massive Text Embedding Benchmark HateSpeechPortugueseClassification is a dataset of Portuguese tweets categorized with their sentiment (2 classes). Task category t2c Domains Social, Written Reference https://aclanthology.org/W19-3510 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HateSpeechPortugueseClassification.texttext-classification1K<n<10K1 likes134 downloads1y agoHugging Face28ccxhwmy /hateful_memes Facebook Hateful Memes Dataset Complete version of the Hateful Memes Challenge dataset (Kiela et al., 2020) with all images included. Dataset Description Hateful memes combine individually benign images and text to produce hateful content. The hate lives in the interaction between modalities, making this one of the hardest content moderation benchmarks. The dataset includes confounders: meme pairs that share the same text (or image) but carry opposite labels, forcing… See the full description on the dataset page: https://huggingface.co/datasets/ccxhwmy/hateful_memes.imageimage-classification10K<n<100K0 likes131 downloads5mo agoHugging Face29panjiyarsunil /hateful-memes-data Hateful Memes (CS5242 submission mirror) Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020) used for reproducibility of our CS5242 (NUS) submission. Contents img/ — 10,000 PNG images of memes train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540), test_seen.jsonl (1,000), test_unseen.jsonl (2,000) Provenance This mirror merges two existing mirrors of the original Meta release: Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/panjiyarsunil/hateful-memes-data.imageimage-classification10K<n<100K0 likes123 downloads19d agoHugging Face30hatemestinbejaia /mmarco-arabic-devtext1M<n<10M0 likes121 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.