datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korea_hate_speechK-MHaS는 추가 레이블링 필수
hate_speech_twitter
Dataset Card for Dataset Name
The dataset is designed to analyze and address hate speech within online platforms. It consists of two sets: the training and testing sets. The two datasets have been labeled and categorized instances of hate speech into nine distinct categories.
Dataset Description
The dataset comprises three key features: tweets, labels (with hate speech denoted as 1 and non-hate speech as 0), and categories (behavior, class, disability, ethnicity, gender… See the full description on the dataset page: https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter.MLMA_hate_speech
Disclaimer
This is a hate speech dataset (in Arabic, French, and English).
Offensive content that does not reflect the opinions of the authors.
Dataset of our EMNLP 2019 Paper (Multilingual and Multi-Aspect Hate Speech Analysis)
For more details about our dataset, please check our paper:
@inproceedings{ousidhoum-etal-multilingual-hate-speech-2019,
title = "Multilingual and Multi-Aspect Hate Speech Analysis",
author = "Ousidhoum, Nedjma… See the full description on the dataset page: https://huggingface.co/datasets/nedjmaou/MLMA_hate_speech.hate_speech_labeledQuran_speech_recognition_kaggleThis dataset can be found in Kaggle
turkish-hate-speech-superset
Turkish Hate Speech Superset
This dataset is a superset (N=41,423) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Turkish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/turkish-hate-speech-superset.american-speech-recognition-dataset
American Speech Dataset for recognition task
Dataset comprises 1,136 hours of telephone dialogues in American, collected from 1,416 native speakers across various topics and domains, achieving an impressive 95% Sentence Accuracy Rate. It is designed for research in automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/american-speech-recognition-dataset.prism-hinglish-hate-speech
PRISM - Code-Mixed Hinglish Hate-Speech Dataset
Binary hate-speech dataset of code-mixed Hindi-English (Hinglish) text, used in the project
Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text
(RSET, The Assam Royal Global University). Source: combined_hate_speech_dataset on Kaggle.
Companion model repository: Hinglish Hate-Speech Classification - BiLSTM / LSTM track
Summary
Attribute
Value
Total samples (raw)
29,550… See the full description on the dataset page: https://huggingface.co/datasets/pankajbiswas6/prism-hinglish-hate-speech.indonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.russian-speech-recognition-dataset
Russian Speech Dataset for recognition task
Dataset comprises 338 hours of telephone dialogues in Russian, collected from 460 native speakers across various topics and domains, with an impressive 98% Word Accuracy Rate. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/russian-speech-recognition-dataset.spanish-hate-speech-superset
Spanish Hate Speech Superset
This dataset is a superset (N=29,855) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Spanish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/spanish-hate-speech-superset.Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_Language2016_2022_hate_speech_filipino
Dataset Card for 2016 and 2022 Hate Speech in Filipino
Dataset Summary
Contains a total of 27,383 tweets that are labeled as hate speech (1) or non-hate speech (0). Split into 80-10-10 (train-validation-test) with a total of 21,773 tweets for training, 2,800 tweets for validation, and 2,810 tweets for testing.
Created by combining hate_speech_filipino and a newly crawled 2022 Philippine Presidential Elections-related Tweets Hate Speech Dataset.
This dataset has an almost… See the full description on the dataset page: https://huggingface.co/datasets/mapsoriano/2016_2022_hate_speech_filipino.slovenian-speech-recognition
Slovenian Speech Dataset
Dataset comprises 10+ hours of audio recordings featuring 20+ speakers engaged in telephone dialogues in the Slovenian language. It contains speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/slovenian-speech-recognition.Dynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.speech_dataset2
NPTEL Telugu Speech Segment Dataset
This dataset contains segment-level metadata and Telugu speech transcriptions generated using Indic Conformer.
Audio Repository
The corresponding complete audio recordings are stored separately in:
Audio Repository – Shravya8-6/speech_dataset
The two datasets are linked using the audio_id field.
For example:
audio_id = telugu_001 in this dataset corresponds to audio_id = telugu_001 in the audio repository.
The corresponding… See the full description on the dataset page: https://huggingface.co/datasets/Shravya8-6/speech_dataset2.Hate-Speech-Tweetsgerman-parliament-speeches
German Parliament Speeches
This dataset contains speeches from the German parliament, derived from the Open Discourse Project (Harvard Dataverse).
Source
Data source:
Open Discourse ProjectHarvard DataverseDOI: 10.7910/DVN/FIKIBO
Original citation:
@data{DVN/FIKIBO_2020,
author = {Richter, Florian and Koch, Philipp and Franke, Oliver and Kraus, Jakob and Kuruc, Fabrizio and Thiem, Anja and Högerl, Judith and Heine, Stella and Schöps, Konstantin},
publisher = {Harvard… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/german-parliament-speeches.vietnamese-speech-recognition
Vietnamese Speech Dataset
Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/vietnamese-speech-recognition.Davidson_Hate_Speech_NewThe_Arabic_News_speech_Corpus_Dataset
Arabic News Speech Corpus Dataset
This dataset is an Arabic speech corpus that supports the development of syllable-based Arabic speech recognition using Wav2Vec-2 architecture and a 5-gram language model. It consists of Modern Standard Arabic (MSA) syllables extracted from TV news broadcasts, annotated with diacritics.
Dataset Details
Dataset Description
This corpus contains 15 hours of WAV audio recordings transcribed into diacritized Modern Standard Arabic… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimSalah/The_Arabic_News_speech_Corpus_Dataset.twitter-hate-speech-en-240ksamplesThis dataset is a combination of the three datasets listed below:
tdavidson/hate_speech_offensive
LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset
ucberkeley-dlab/measuring-hate-speech
It has only two columns, "tweet" and "labels", and 242738 rows of uncleaned data.
Chinese-Speech-Dataset
🎧 Chinese (Simplified) Speech Dataset
The Chinese (Simplified) speech dataset is a high-quality speech audio dataset developed to support scalable AI and machine learning solutions with diverse and structured audio data. It contains 105 hours of speech data across 700 audio files, delivered in MP3 and WAV formats, with a total size of 229 MB. This well-balanced audio dataset provides reliable voice data, featuring 54% female and 46% male speakers, with age distribution ranging from… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Chinese-Speech-Dataset.japanese-speech-recognition-dataset
Japanese Speech Dataset for recognition task
Dataset comprises 10+ hours of telephone dialogues in Japanese, collected from 10 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/japanese-speech-recognition-dataset.apple-speechanalyzer-vs-whisper-cpp-mac
Apple SpeechAnalyzer vs whisper.cpp on Mac
Four complete speech-recognition benchmark runs over the same deterministic
40-speaker LibriSpeech test-clean snapshot:
Engine
Model path
WER
CER
Repeated median post-speech latency
Repeated p95
Apple SpeechAnalyzer
progressiveTranscription on macOS 26.5
1.98%
1.02%
125–132 ms
194–201 ms
whisper.cpp server
1.8.4 · ggml-small.en
4.28%
1.79%
122–125 ms
152–161 ms
Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.gender-hate-speechThe "gender identity" subset of the large-scale dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This subset is used in the experiments of "Şahinuç, F., Yilmaz, E. H., Toraman, C., & Koç, A. (2023). The effect of gender bias on hate speech detection. Signal, Image and Video Processing, 17(4), 1591-1597."
The "gender identity" subset includes 20,000 tweets in English.
The published data split is the first fold of 10-fold cross-validation… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/gender-hate-speech.Deepfake-Eval-2024-Protocals
WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection
Download the Protocols
Install the datasets package:
pip install datasets
Log in with your Hugging Face account:
huggingface-cli login
Load the dataset in Python:
from datasets import load_dataset
# Download from HF and cache
ds = load_dataset("xxuan-speech/Deepfake-Eval-2024-Protocals")
Statistics of Deepfake-Eval-2024 Benchmark
Dataset
Total
Real
Fake… See the full description on the dataset page: https://huggingface.co/datasets/xxuan-speech/Deepfake-Eval-2024-Protocals.arabic-hate-speech-superset
Arabic Hate Speech Superset
This dataset is a superset (N=449,078) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Arabic hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/arabic-hate-speech-superset.Yemeni-Speech-Emotion-Dataset
YSED — Yemeni Speech Emotion Dataset (audio-classification repackaging)
A clean repackaging of YSED with a metadata.csv and stratified train/validation/test splits, for emotion classification on Yemeni Arabic.
Original dataset: Derhem, S., AL-Mekhlafi, E., AL-Majmar, N. A., & AL-Makhlafi, M. (2025). YSED: Yemeni Speech Emotion Dataset. Data in Brief. DOI: 10.1016/j.dib.2025.112233. Zenodo: https://zenodo.org/records/15227219.
What's in here
1432 audio clips across… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Yemeni-Speech-Emotion-Dataset.Korean-Speech-Dataset
🎧 Korean Speech Dataset
The Korean Speech Dataset is a large-scale speech audio dataset designed to provide high-quality and structured audio data for advanced AI and machine learning systems. It includes 192 hours of audio data across 628 files, delivered in MP3 and WAV formats, with a total size of 447 MB. This well-balanced audio dataset ensures diverse and representative voice data, with 52% female and 48% male speakers, and an age distribution ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Korean-Speech-Dataset.
