datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hate-Speech-Tweetsturkish-hate-speech-superset
Turkish Hate Speech Superset
This dataset is a superset (N=41,423) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Turkish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/turkish-hate-speech-superset.Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_LanguageDynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.spanish-hate-speech-superset
Spanish Hate Speech Superset
This dataset is a superset (N=29,855) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Spanish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/spanish-hate-speech-superset.speech_dataset2
NPTEL Telugu Speech Segment Dataset
This dataset contains segment-level metadata and Telugu speech transcriptions generated using Indic Conformer.
Audio Repository
The corresponding complete audio recordings are stored separately in:
Audio Repository – Shravya8-6/speech_dataset
The two datasets are linked using the audio_id field.
For example:
audio_id = telugu_001 in this dataset corresponds to audio_id = telugu_001 in the audio repository.
The corresponding… See the full description on the dataset page: https://huggingface.co/datasets/Shravya8-6/speech_dataset2.german-parliament-speeches
German Parliament Speeches
This dataset contains speeches from the German parliament, derived from the Open Discourse Project (Harvard Dataverse).
Source
Data source:
Open Discourse ProjectHarvard DataverseDOI: 10.7910/DVN/FIKIBO
Original citation:
@data{DVN/FIKIBO_2020,
author = {Richter, Florian and Koch, Philipp and Franke, Oliver and Kraus, Jakob and Kuruc, Fabrizio and Thiem, Anja and Högerl, Judith and Heine, Stella and Schöps, Konstantin},
publisher = {Harvard… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/german-parliament-speeches.welsh-speech-3d-meshes
Welsh Speech Dataset - 3D Facial Meshes
3D facial reconstructions from the Welsh Speech Dataset.
Contents
3D meshes (.obj files) - One per frame
Texture maps (.png files) - Fused left-right stereo images from 3DMD
Captured using 3DMD 6-camera system
~330 zip files (one per speaker-phrase sequence)
File Structure
Files are organized as zip archives in the meshes/ directory, one zip per speaker-phrase sequence:
meshes/
├── speaker_01_phrase_01.zip
├──… See the full description on the dataset page: https://huggingface.co/datasets/arvinsingh/welsh-speech-3d-meshes.arabic-hate-speech-superset
Arabic Hate Speech Superset
This dataset is a superset (N=449,078) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Arabic hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/arabic-hate-speech-superset.apple-speechanalyzer-vs-whisper-cpp-mac
Apple SpeechAnalyzer vs whisper.cpp on Mac
Four complete speech-recognition benchmark runs over the same deterministic
40-speaker LibriSpeech test-clean snapshot:
Engine
Model path
WER
CER
Repeated median post-speech latency
Repeated p95
Apple SpeechAnalyzer
progressiveTranscription on macOS 26.5
1.98%
1.02%
125–132 ms
194–201 ms
whisper.cpp server
1.8.4 · ggml-small.en
4.28%
1.79%
122–125 ms
152–161 ms
Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.Yemeni-Speech-Emotion-Dataset
YSED — Yemeni Speech Emotion Dataset (audio-classification repackaging)
A clean repackaging of YSED with a metadata.csv and stratified train/validation/test splits, for emotion classification on Yemeni Arabic.
Original dataset: Derhem, S., AL-Mekhlafi, E., AL-Majmar, N. A., & AL-Makhlafi, M. (2025). YSED: Yemeni Speech Emotion Dataset. Data in Brief. DOI: 10.1016/j.dib.2025.112233. Zenodo: https://zenodo.org/records/15227219.
What's in here
1432 audio clips across… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Yemeni-Speech-Emotion-Dataset.VN-SpeechMix_Dataset
VN-SpeechMix: A Large-Scale Multi-Dialect Vietnamese Speech Mixture Dataset
VN-SpeechMix is a large-scale, multi-dialect Vietnamese speech mixture
dataset for two-speaker speech separation research. It is built from the
ViMD corpus (Van Dinh et al., EMNLP 2024)
using a loudness-aware mixing pipeline (LUFS normalization + two-stage
anti-clipping) and a dialect-aware pairing strategy across Vietnam's three
macro-dialect regions (North / Central / South).
26,000 two-speaker… See the full description on the dataset page: https://huggingface.co/datasets/pervasiveaidataresearchlab2025/VN-SpeechMix_Dataset.large-scale-hate-speech-turkish-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1 (Turkish):
The original dataset that includes 100,000 tweets in Turkish. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v1.large-scale-hate-speech-turkish-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2 (Turkish):
The modified dataset that includes 60,310 tweets in Turkish. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v2.tech-speech-congress
Tech-Speech in the Congressional Record, 1995–2025
Descriptive aggregates from a speech-level analysis of technology-related
discourse in the U.S. Congressional Record, 1995–2025. Built from
govinfo.gov daily Congressional Record packages, parsed to the individual
speech turn with speaker metadata (party, state, chamber, Bioguide ID),
filtered for procedural speech, and scored with a TF-IDF technology-intensity
index over a 237-term technology vocabulary built from federal and… See the full description on the dataset page: https://huggingface.co/datasets/tapanyemre/tech-speech-congress.lost-in-speech
Lost in Speech
A trilingual benchmark for reference-free classification of synthetically introduced factual and contextual alterations in English, Russian, and Kazakh. It contains 12,013 samples derived from news articles, with text, synthesized speech, and ASR transcript representations used in the study.
Altered samples are LLM-generated rewrites with a controlled alteration type—contradiction, fabrication, or context inconsistency—and severity level—mild, moderate, or severe.… See the full description on the dataset page: https://huggingface.co/datasets/maristombayeva/lost-in-speech.multilabel-tagalog-hate-speechfrench-hate-speech-superset
French Hate Speech Superset
This dataset is a superset (N=18,071) of posts annotated as hateful or not. It results from the preprocessing and merge of all available French hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior, that… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/french-hate-speech-superset.gigaspeech-testlarge-scale-hate-speech-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1:
The original dataset that includes 100,000 tweets in English. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 1 (English)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:
NOTE:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-v1.german-hate-speech-superset
German Hate Speech Superset
This dataset is a superset (N=50,545) of posts annotated as hateful or not. It results from the preprocessing and merge of all available German hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior, that… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/german-hate-speech-superset.hate_speech_open_data_original_class_test_setindonesian-hate-speech-superset
Indonesian Hate Speech Superset
This dataset is a superset (N=14,306) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Indonesian hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/indonesian-hate-speech-superset.portuguese-hate-speech-superset
Portuguese Hate Speech Superset
This dataset is a superset (N=43,222) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Portuguese hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/portuguese-hate-speech-superset.hate-speech-targethttps://coltekin.github.io/offensive-turkish/guidelines-tr.html
large-scale-hate-speech-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2:
The modified dataset that includes 68,597 tweets in English. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 1 (English)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:
NOTE:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-v2.Multi-Label_Bangla_Hate_Speech_Datareadme_text = """
Bangla Hate Speech Extended Dataset
📖 Overview
This dataset is an expanded version of the original Bengali Hate Speech Dataset created by Hriteshwar Talukder and Md Saiful Islam.
The original dataset provided a strong foundation for hate speech detection in the Bengali language. In this extended version, the dataset has been:
Expanded in size with ~5000 additional Bengali social media comments.
Reclassified with fine-grained categories… See the full description on the dataset page: https://huggingface.co/datasets/sumaiya-afroze/Multi-Label_Bangla_Hate_Speech_Data.revised_Toraman22_hate_speech_v2
Dataset Card for Dataset Name
This dataset card is the revised and cleaned dataset introduced by Toraman et al., LREC 2022.
Dataset Details
It contains a total of 68590 English tweets ready for processing.
Each tweet has either of the three labels 0 - Normal, 1 - Offensive and 2 - Hate
Each tweet has either of the five domains 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
Dataset contains 68,590 English tweets ready for processing.
Each tweet is categorized with… See the full description on the dataset page: https://huggingface.co/datasets/mahmed31/revised_Toraman22_hate_speech_v2.SpeechPatternDisorderData
SpeechPatternDisorderData
tags: speech pattern recognition, speech disorder diagnosis, clinical application
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'SpeechPatternDisorderData' CSV dataset is designed to support machine learning models in the recognition and diagnosis of speech disorders through pattern analysis. It includes a diverse collection of audio clips and corresponding text transcriptions from various speakers… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/SpeechPatternDisorderData.meps_speeches_with_translation.csv
