CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SihyunPark /korea_hate_speechK-MHaS는 추가 레이블링 필수 text100K<n<1M0 likes354 downloads2y agoHugging Face02thefrankhsu /hate_speech_twitter Dataset Card for Dataset Name The dataset is designed to analyze and address hate speech within online platforms. It consists of two sets: the training and testing sets. The two datasets have been labeled and categorized instances of hate speech into nine distinct categories. Dataset Description The dataset comprises three key features: tweets, labels (with hate speech denoted as 1 and non-hate speech as 0), and categories (behavior, class, disability, ethnicity, gender… See the full description on the dataset page: https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter.texttext-classification1K<n<10K5 likes327 downloads3y agoHugging Face03nedjmaou /MLMA_hate_speech Disclaimer This is a hate speech dataset (in Arabic, French, and English). Offensive content that does not reflect the opinions of the authors. Dataset of our EMNLP 2019 Paper (Multilingual and Multi-Aspect Hate Speech Analysis) For more details about our dataset, please check our paper: @inproceedings{ousidhoum-etal-multilingual-hate-speech-2019, title = "Multilingual and Multi-Aspect Hate Speech Analysis", author = "Ousidhoum, Nedjma… See the full description on the dataset page: https://huggingface.co/datasets/nedjmaou/MLMA_hate_speech.text10K<n<100K5 likes215 downloads2y agoHugging Face04Doowon96 /hate_speech_labeledtext1K<n<10K0 likes199 downloads3y agoHugging Face05Nuwaisir /Quran_speech_recognition_kaggleThis dataset can be found in Kaggle text10K<n<100K4 likes170 downloads5y agoHugging Face06manueltonneau /turkish-hate-speech-supersetgated Turkish Hate Speech Superset This dataset is a superset (N=41,423) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Turkish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/turkish-hate-speech-superset.tabulartext-classification10K<n<100K2 likes131 downloads2y agoHugging Face07UniDataPro /american-speech-recognition-dataset American Speech Dataset for recognition task Dataset comprises 1,136 hours of telephone dialogues in American, collected from 1,416 native speakers across various topics and domains, achieving an impressive 95% Sentence Accuracy Rate. It is designed for research in automatic speech recognition (ASR) systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/american-speech-recognition-dataset.textn<1K5 likes124 downloads1mo agoHugging Face08pankajbiswas6 /prism-hinglish-hate-speech PRISM - Code-Mixed Hinglish Hate-Speech Dataset Binary hate-speech dataset of code-mixed Hindi-English (Hinglish) text, used in the project Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text (RSET, The Assam Royal Global University). Source: combined_hate_speech_dataset on Kaggle. Companion model repository: Hinglish Hate-Speech Classification - BiLSTM / LSTM track Summary Attribute Value Total samples (raw) 29,550… See the full description on the dataset page: https://huggingface.co/datasets/pankajbiswas6/prism-hinglish-hate-speech.texttext-classification10K<n<100K0 likes123 downloads3mo agoHugging Face09haipradana /indonesian-twitter-hate-speech-cleaned Dataset Card for indonesian-twitter-hate-speech-cleaned Dataset Summary Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories. The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.texttext-classification10K<n<100K0 likes118 downloads1y agoHugging Face10UniDataPro /russian-speech-recognition-dataset Russian Speech Dataset for recognition task Dataset comprises 338 hours of telephone dialogues in Russian, collected from 460 native speakers across various topics and domains, with an impressive 98% Word Accuracy Rate. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. By utilizing this dataset, researchers and developers can advance their… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/russian-speech-recognition-dataset.textn<1K7 likes99 downloads1mo agoHugging Face11manueltonneau /spanish-hate-speech-supersetgated Spanish Hate Speech Superset This dataset is a superset (N=29,855) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Spanish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available or could be retrieved with the Twitter API focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/spanish-hate-speech-superset.tabulartext-classification10K<n<100K6 likes96 downloads2y agoHugging Face12dirtycomputer /Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_Languagetabular10K<n<100K0 likes89 downloads3y agoHugging Face13mapsoriano /2016_2022_hate_speech_filipino Dataset Card for 2016 and 2022 Hate Speech in Filipino Dataset Summary Contains a total of 27,383 tweets that are labeled as hate speech (1) or non-hate speech (0). Split into 80-10-10 (train-validation-test) with a total of 21,773 tweets for training, 2,800 tweets for validation, and 2,810 tweets for testing. Created by combining hate_speech_filipino and a newly crawled 2022 Philippine Presidential Elections-related Tweets Hate Speech Dataset. This dataset has an almost… See the full description on the dataset page: https://huggingface.co/datasets/mapsoriano/2016_2022_hate_speech_filipino.texttext-classification10K<n<100K1 likes87 downloads2y agoHugging Face14UniDataPro /slovenian-speech-recognition Slovenian Speech Dataset Dataset comprises 10+ hours of audio recordings featuring 20+ speakers engaged in telephone dialogues in the Slovenian language. It contains speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/slovenian-speech-recognition.audioautomatic-speech-recognitionn<1K2 likes86 downloads1mo agoHugging Face15LennardZuendorf /Dynamically-Generated-Hate-Speech-Dataset Dataset Card for dynamically generated hate speech dataset Dataset Summary This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela Original README from GitHub Dynamically-Generated-Hate-Speech-Dataset ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.tabulartext-classification10K<n<100K6 likes83 downloads3y agoHugging Face16Shravya8-6 /speech_dataset2 NPTEL Telugu Speech Segment Dataset This dataset contains segment-level metadata and Telugu speech transcriptions generated using Indic Conformer. Audio Repository The corresponding complete audio recordings are stored separately in: Audio Repository – Shravya8-6/speech_dataset The two datasets are linked using the audio_id field. For example: audio_id = telugu_001 in this dataset corresponds to audio_id = telugu_001 in the audio repository. The corresponding… See the full description on the dataset page: https://huggingface.co/datasets/Shravya8-6/speech_dataset2.tabular10K<n<100K0 likes79 downloads26d agoHugging Face17kaifahmad /Hate-Speech-Tweetstabular10K<n<100K0 likes76 downloads3y agoHugging Face18emilpartow /german-parliament-speeches German Parliament Speeches This dataset contains speeches from the German parliament, derived from the Open Discourse Project (Harvard Dataverse). Source Data source: Open Discourse ProjectHarvard DataverseDOI: 10.7910/DVN/FIKIBO Original citation: @data{DVN/FIKIBO_2020, author = {Richter, Florian and Koch, Philipp and Franke, Oliver and Kraus, Jakob and Kuruc, Fabrizio and Thiem, Anja and Högerl, Judith and Heine, Stella and Schöps, Konstantin}, publisher = {Harvard… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/german-parliament-speeches.tabulartext-classification100K<n<1M4 likes72 downloads1y agoHugging Face19UniDataPro /vietnamese-speech-recognition Vietnamese Speech Dataset Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/vietnamese-speech-recognition.audioautomatic-speech-recognitionn<1K4 likes71 downloads1mo agoHugging Face20krishan-CSE /Davidson_Hate_Speech_Newtext10K<n<100K0 likes69 downloads3y agoHugging Face21IbrahimSalah /The_Arabic_News_speech_Corpus_Dataset Arabic News Speech Corpus Dataset This dataset is an Arabic speech corpus that supports the development of syllable-based Arabic speech recognition using Wav2Vec-2 architecture and a 5-gram language model. It consists of Modern Standard Arabic (MSA) syllables extracted from TV news broadcasts, annotated with diacritics. Dataset Details Dataset Description This corpus contains 15 hours of WAV audio recordings transcribed into diacritized Modern Standard Arabic… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimSalah/The_Arabic_News_speech_Corpus_Dataset.audioautomatic-speech-recognition1K<n<10K6 likes69 downloads2y agoHugging Face22TLeonidas /twitter-hate-speech-en-240ksamplesThis dataset is a combination of the three datasets listed below: tdavidson/hate_speech_offensive LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset ucberkeley-dlab/measuring-hate-speech It has only two columns, "tweet" and "labels", and 242738 rows of uncleaned data. text100K<n<1M1 likes67 downloads2y agoHugging Face23Speech-data /Chinese-Speech-Dataset 🎧 Chinese (Simplified) Speech Dataset The Chinese (Simplified) speech dataset is a high-quality speech audio dataset developed to support scalable AI and machine learning solutions with diverse and structured audio data. It contains 105 hours of speech data across 700 audio files, delivered in MP3 and WAV formats, with a total size of 229 MB. This well-balanced audio dataset provides reliable voice data, featuring 54% female and 46% male speakers, with age distribution ranging from… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Chinese-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes63 downloads6mo agoHugging Face24UniDataPro /japanese-speech-recognition-dataset Japanese Speech Dataset for recognition task Dataset comprises 10+ hours of telephone dialogues in Japanese, collected from 10 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/japanese-speech-recognition-dataset.audion<1K1 likes60 downloads1mo agoHugging Face25researchaudio /apple-speechanalyzer-vs-whisper-cpp-mac Apple SpeechAnalyzer vs whisper.cpp on Mac Four complete speech-recognition benchmark runs over the same deterministic 40-speaker LibriSpeech test-clean snapshot: Engine Model path WER CER Repeated median post-speech latency Repeated p95 Apple SpeechAnalyzer progressiveTranscription on macOS 26.5 1.98% 1.02% 125–132 ms 194–201 ms whisper.cpp server 1.8.4 · ggml-small.en 4.28% 1.79% 122–125 ms 152–161 ms Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.tabularautomatic-speech-recognitionn<1K0 likes55 downloads2mo agoHugging Face26ctoraman /gender-hate-speechThe "gender identity" subset of the large-scale dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer". This subset is used in the experiments of "Şahinuç, F., Yilmaz, E. H., Toraman, C., & Koç, A. (2023). The effect of gender bias on hate speech detection. Signal, Image and Video Processing, 17(4), 1591-1597." The "gender identity" subset includes 20,000 tweets in English. The published data split is the first fold of 10-fold cross-validation… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/gender-hate-speech.texttext-classification10K<n<100K3 likes54 downloads3y agoHugging Face27xxuan-speech /Deepfake-Eval-2024-Protocals WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection Download the Protocols Install the datasets package: pip install datasets Log in with your Hugging Face account: huggingface-cli login Load the dataset in Python: from datasets import load_dataset # Download from HF and cache ds = load_dataset("xxuan-speech/Deepfake-Eval-2024-Protocals") Statistics of Deepfake-Eval-2024 Benchmark Dataset Total Real Fake… See the full description on the dataset page: https://huggingface.co/datasets/xxuan-speech/Deepfake-Eval-2024-Protocals.text10K<n<100K0 likes54 downloads8mo agoHugging Face28manueltonneau /arabic-hate-speech-supersetgated Arabic Hate Speech Superset This dataset is a superset (N=449,078) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Arabic hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available or could be retrieved with the Twitter API focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/arabic-hate-speech-superset.tabulartext-classification100K<n<1M8 likes53 downloads2y agoHugging Face29FatimahEmadEldin /Yemeni-Speech-Emotion-Dataset YSED — Yemeni Speech Emotion Dataset (audio-classification repackaging) A clean repackaging of YSED with a metadata.csv and stratified train/validation/test splits, for emotion classification on Yemeni Arabic. Original dataset: Derhem, S., AL-Mekhlafi, E., AL-Majmar, N. A., & AL-Makhlafi, M. (2025). YSED: Yemeni Speech Emotion Dataset. Data in Brief. DOI: 10.1016/j.dib.2025.112233. Zenodo: https://zenodo.org/records/15227219. What's in here 1432 audio clips across… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Yemeni-Speech-Emotion-Dataset.audioaudio-classification1K<n<10K1 likes52 downloads5mo agoHugging Face30Speech-data /Korean-Speech-Dataset 🎧 Korean Speech Dataset The Korean Speech Dataset is a large-scale speech audio dataset designed to provide high-quality and structured audio data for advanced AI and machine learning systems. It includes 192 hours of audio data across 628 files, delivered in MP3 and WAV formats, with a total size of 447 MB. This well-balanced audio dataset ensures diverse and representative voice data, with 52% female and 48% male speakers, and an age distribution ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Korean-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes48 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.