datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PMC
Data collected from PMC
Only CC-BY, CC-BY-SA licenses are included.
For all records, check the jsonl files in the data folder
korean-hate-speechreference: https://github.com/kocohub/korean-hate-speech
@inproceedings{moon-etal-2020-beep,
title = "{BEEP}! {K}orean Corpus of Online News Comments for Toxic Speech Detection",
author = "Moon, Jihyung and
Cho, Won Ik and
Lee, Junbum",
booktitle = "Proceedings of the Eighth International Workshop on Natural Language Processing for Social Media",
month = jul,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics"… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/korean-hate-speech.measuring-hate-speech
Dataset card for Measuring Hate Speech
This is a public release of the dataset described in Kennedy et al. (2020) and Sachdeva et al. (2022), consisting of 39,565 comments annotated by 7,912 annotators, for 135,556 combined rows. The primary outcome variable is the "hate speech score" but the 10 constituent ordinal labels (sentiment, (dis)respect, insult, humiliation, inferior status, violence, dehumanization, genocide, attack/defense, hate speech benchmark) can also be treated as… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/measuring-hate-speech.hate_speech_offensive
Dataset Card for [Dataset Name]
Dataset Summary
An annotated dataset for hate speech and offensive language detection on tweets.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English (en)
Dataset Structure
Data Instances
{
"count": 3,
"hate_speech_annotation": 0,
"offensive_language_annotation": 0,
"neither_annotation": 3,
"label": 2, # "neither"
"tweet": "!!! RT @mayasolovely: As a woman you… See the full description on the dataset page: https://huggingface.co/datasets/tdavidson/hate_speech_offensive.White-Hat-Security-Agent-Prompts-600K
White Hat Security Agent Prompts 600K
Overview
The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios.
Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.hateful-memes-data
Hateful Memes (CS5242 submission mirror)
Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020)
used for reproducibility of our CS5242 (NUS) submission.
Contents
img/ — 10,000 PNG images of memes
train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540),
test_seen.jsonl (1,000), test_unseen.jsonl (2,000)
Provenance
This mirror merges two existing mirrors of the original Meta release:
Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/cs5242-hateful-memes/hateful-memes-data.tweets_hate_speech_detection
Dataset Card for Tweets Hate Speech Detection
Dataset Summary
The objective of this task is to detect hate speech in tweets. For the sake of simplicity, we say a tweet contains hate speech if it has a racist or sexist sentiment associated with it. So, the task is to classify racist or sexist tweets from other tweets.
Formally, given a training sample of tweets and labels, where label ‘1’ denotes the tweet is racist/sexist and label ‘0’ denotes the tweet is not… See the full description on the dataset page: https://huggingface.co/datasets/tweets-hate-speech-detection/tweets_hate_speech_detection.ndlj_tosho_1
国会図書館に収蔵される著作権切れのデータです
galaxy-mentions-hats
Galaxy Mentions HATS
A HATS catalog of 43,546 resolved literature mentions from
astronolan/galaxy-mentions,
prepared for efficient spatial crossmatching with the
Multimodal Universe HATS catalogs.
Each row is a mention in a paper, not a deduplicated astronomical object. Multiple rows may
therefore describe the same galaxy. The optional wiki_entity_id groups mentions using the current
1 arcsecond connected-components build, while mention_id remains the unique row identifier.… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/galaxy-mentions-hats.miracl-arabicjapanese2010
日本語ウェブコーパス2010
こちらのデータをhuggingfaceにアップロードしたものです。
2009 年度における著作権法の改正(平成21年通常国会 著作権法改正等について | 文化庁)に基づき,情報解析研究への利用に限って利用可能です。
形態素解析を用いて、自動で句点をつけました。
変換コード
変換スクリプト
形態素解析など
CommonCrawl_wet_v2roman_urdu_hate_speech
Dataset Card for roman_urdu_hate_speech
Dataset Summary
The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.synth_hathor_20000hat
HAT: Hallucination Annotation for Translation
🧭 Table of Contents
Overview
Usage
Data Creation Process
Data Statistics
Dataset Structure
Paper Abstract
Citation
License
📘 Overview
HAT (Hallucination Annotation for Translation) is a large-scale dataset for hallucination detection in machine translation (MT).It is released as part of our publication at ACL 2026 (paper).
350,959 span-level annotated samples
38 language pairs
~8,000-10,000… See the full description on the dataset page: https://huggingface.co/datasets/apple/hat.MMSoc_HatefulMemesbn_hate_speech
Dataset Card for Bengali Hate Speech Dataset
Dataset Summary
The Bengali Hate Speech Dataset is a Bengali-language dataset of news articles collected from various Bengali media sources and categorized based on the type of hate in the text. The dataset was created to provide greater support for under-resourced languages like Bengali on NLP tasks, and serves as a benchmark for multiple types of classification tasks.
Supported Tasks and Leaderboards
topic… See the full description on the dataset page: https://huggingface.co/datasets/rezacsedu/bn_hate_speech.STCALIR_Synthetic-Test-Collectionhateful_memes
Facebook Hateful Memes Dataset
Complete version of the Hateful Memes Challenge
dataset (Kiela et al., 2020) with all images included.
Dataset Description
Hateful memes combine individually benign images and text to produce hateful
content. The hate lives in the interaction between modalities, making this
one of the hardest content moderation benchmarks.
The dataset includes confounders: meme pairs that share the same text (or
image) but carry opposite labels, forcing… See the full description on the dataset page: https://huggingface.co/datasets/dffeewew/hateful_memes.contextualized_hate_speech
Contextualized Hate Speech: A dataset of comments in news outlets on Twitter
Dataset Summary
This dataset is a collection of tweets that were posted in response to news articles from five specific Argentinean news outlets: Clarín, Infobae, La Nación, Perfil and Crónica, during the COVID-19 pandemic. The comments were analyzed for hate speech across eight different characteristics: against women, racist content, class hatred, against LGBTQ+ individuals, against physical… See the full description on the dataset page: https://huggingface.co/datasets/piuba-bigdata/contextualized_hate_speech.vqa_hataw_v1codemixed-id-hate-speech
Code-mixed Indonesian Hate Speech Dataset
Manually annotated hate speech dataset for Indonesian-Javanese and
Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations.
task905_hate_speech_offensive_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task905_hate_speech_offensive_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task905_hate_speech_offensive_classification.hateXplain_filteredhate_speech_pl
Dataset Card for HateSpeechPl
Dataset Summary
The dataset was created to analyze the possibility of automating the recognition of hate speech in Polish. It was collected from the Polish forums and represents various types and degrees of offensive language, expressed towards minorities.
The original dataset is provided as an export of MySQL tables, what makes it hard to load. Due to that, it was converted to CSV and put to a Github repository.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hate_speech_pl.hatespeech-ind-multilabelclassification
HateSpeech_ind_MultiLabelClassification
Deduplicated copy of kornwtp/hatespeech-ind-multilabelclassification.
Splits
split
rows
train
13,014
HateSpeechPortugueseClassification
HateSpeechPortugueseClassification
An MTEB dataset
Massive Text Embedding Benchmark
HateSpeechPortugueseClassification is a dataset of Portuguese tweets categorized with their sentiment (2 classes).
Task category
t2c
Domains
Social, Written
Reference
https://aclanthology.org/W19-3510
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HateSpeechPortugueseClassification.hateful_memes
Facebook Hateful Memes Dataset
Complete version of the Hateful Memes Challenge
dataset (Kiela et al., 2020) with all images included.
Dataset Description
Hateful memes combine individually benign images and text to produce hateful
content. The hate lives in the interaction between modalities, making this
one of the hardest content moderation benchmarks.
The dataset includes confounders: meme pairs that share the same text (or
image) but carry opposite labels, forcing… See the full description on the dataset page: https://huggingface.co/datasets/ccxhwmy/hateful_memes.hateful-memes-data
Hateful Memes (CS5242 submission mirror)
Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020)
used for reproducibility of our CS5242 (NUS) submission.
Contents
img/ — 10,000 PNG images of memes
train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540),
test_seen.jsonl (1,000), test_unseen.jsonl (2,000)
Provenance
This mirror merges two existing mirrors of the original Meta release:
Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/panjiyarsunil/hateful-memes-data.mmarco-arabic-dev
