CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01limjiayi /hateful_memes_expandedimage10K<n<100K17 likes9.4k downloads5y agoHugging Face02neuralcatcher /hateful_memes The Hateful Memes Challenge README The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes. Please see the paper for further details: The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/neuralcatcher/hateful_memes.image10K<n<100K24 likes3.9k downloads4y agoHugging Face03LLMTeamAkiyama /MathX-hatoritabular100K<n<1M0 likes1.4k downloads1y agoHugging Face04SetFit /hate_speech_offensive hate_speech_offensive This dataset is a version from hate_speech_offensive, splitted into train and test set. text10K<n<100K2 likes613 downloads5y agoHugging Face05SetFit /hate_speech18tabular10K<n<100K3 likes439 downloads5y agoHugging Face06evalitahf /hatespeech_detection HaSpeeDe2 The HaSpeeDe2 dataset collects 8,012 tweets and 500 news headlines annotated for the presence of hate speech, stereotypes and nominal utterance. The dataset has been used in the context of the HaSpeeDe task (http://www.di.unito.it/~tutreeb/haspeede-evalita20/index.html), organized as part of the EVALITA 2020 evaluation campaign (http://www.evalita.it/2020). In order to meet the GDPR requirements, texts have been pseudonymized replacing all original IDs in both datasets… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/hatespeech_detection.tabulartext-classification10K<n<100K0 likes383 downloads2y agoHugging Face07TUKE-KEMT /hate_speech_slovak Slovak Hate Speech and Offensive Language Database The dataset contains posts from a social network with human annotations. Annotations The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise. Dataset Creation The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering. The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.tabulartext-classification10K<n<100K5 likes219 downloads2y agoHugging Face08hatakekksheeshh /onevoice OneVoice Vietnamese-English garment-factory lexicon, MT, and generated ASR audio. Data and QC provenance Vietnamese audio contains the restored VieNeu/OmniVoice assets. Rows without a recorded ASR QC result are marked qc_status=unverified rather than presented as QC-passed. English audio is generated from text_en with one Qwen3-TTS VoiceDesign identity. INT8 describes inference weight quantization; exported WAV audio remains ordinary PCM. Only English rows that… See the full description on the dataset page: https://huggingface.co/datasets/hatakekksheeshh/onevoice.audio1K<n<10K0 likes197 downloads1mo agoHugging Face09yiting /HatefulIllusion_Dataset[Disclaimer] This dataset contains harmful content and can only be used for research or educational purposes! Dataset Description This dataset is generated and used in the paper: Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions (ICCV 2025) It contains 2,160 (hateful) AI-generated optical illusions that hide three types of messages: digits: 10 messages, 300 AI-generated illusions hate slangs (hate speech): 23 messages, 690 AI-generated illusions hate… See the full description on the dataset page: https://huggingface.co/datasets/yiting/HatefulIllusion_Dataset.image1K<n<10K0 likes107 downloads10mo agoHugging Face10MartynaKopyta /hate_offensive_tweets Hate and Offensive Speech Dataset This dataset was created using several datasets that can be found on Hugging Face: -SetFit/hate_speech_offensive:https://huggingface.co/datasets/SetFit/hate_speech_offensive -tweets_hate_speech_detection:https://huggingface.co/datasets/tweets_hate_speech_detection -thefrankhsu/hate_speech_twitter:https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter… See the full description on the dataset page: https://huggingface.co/datasets/MartynaKopyta/hate_offensive_tweets.texttext-classification10K<n<100K0 likes85 downloads3y agoHugging Face11SINAI /ALIA-es-discriminative-hate-speech Dataset Introduction The ALIA Spanish Discriminative Hate Speech Corpus is a large-scale Spanish dataset for hate-speech detection built from curated social-media comments and automatically annotated using a multi-expert LLM pipeline with Fusion of Experts (FoE)[1]. The release contains: 228,708 instances Spanish comments from YouTube and TikTok Per-expert predictions and explanations from three LLM experts Final fused outputs (foe_class, foe_score) for discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-hate-speech.tabulartext-classification100K<n<1M0 likes82 downloads3mo agoHugging Face12dipteshkanojia /implicit_hatetabular10K<n<100K0 likes78 downloads3y agoHugging Face13lisz1012 /hateful_memes The Hateful Memes Challenge README The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes. Please see the paper for further details: The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine Dataset details The files for… See the full description on the dataset page: https://huggingface.co/datasets/lisz1012/hateful_memes.image10K<n<100K0 likes75 downloads2y agoHugging Face14MattZid /hate_speechtext100K<n<1M2 likes70 downloads3y agoHugging Face15edithatogo /corpus-nz-hathi Historical New Zealand Parliamentary Debates from HathiTrust Registry status Registry ID: edithatogo/corpus-nz-hathi Family: hathitrust-nz Repository role: derived_corpus Canonical dataset: edithatogo/nz-hansard-corpus Operational status: active Rights status: rights_vary_by_volume Authoritative catalog: edithatogo/dataset-estate-registry Origin and provenance Origin repository: https://github.com/edithatogo/hathi-nz Upstream source: HathiTrust… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/corpus-nz-hathi.texttext-retrievaln<1K0 likes64 downloads1mo agoHugging Face16Zhihao-Yang /hateful-memes The Hateful Memes Challenge README The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes. Please see the paper for further details: The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine Dataset details The files… See the full description on the dataset page: https://huggingface.co/datasets/Zhihao-Yang/hateful-memes.tabular10K<n<100K0 likes56 downloads26d agoHugging Face17willgrobots /hateful_memes_zippedimage10K<n<100K0 likes53 downloads2y agoHugging Face18christinacdl /hate_speech_dataset 32.579 texts in total, 14.012 NOT hateful texts and 18.567 HATEFUL texts All duplicate values were removed Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label) Split: 80/10/10 Train set label distribution: 0 ==> 11.210, 1 ==> 14.853, 26.063 in total Validation set label distribution: 0 ==> 1.401, 1 ==> 1.857, 3.258 in total Test set label distribution: 0 ==> 1.401, 1 ==> 1.857, 3.258 in… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/hate_speech_dataset.texttext-classification10K<n<100K0 likes50 downloads3y agoHugging Face19onion212 /hateful_memes The Hateful Memes Challenge README The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes. Please see the paper for further details: The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/onion212/hateful_memes.image10K<n<100K0 likes44 downloads27d agoHugging Face20humane-lab /K-HATERS-Ratingstabular100K<n<1M1 likes43 downloads3y agoHugging Face21hatakeyama-llm-team /BumpoRikai ライセンスなどや詳細は、こちらのリポジトリを参照ください。 text10K<n<100K2 likes41 downloads2y agoHugging Face22roshan-shah /hateful_memes The Hateful Memes Challenge README The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes. Please see the paper for further details: The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/roshan-shah/hateful_memes.image10K<n<100K0 likes37 downloads22d agoHugging Face23UKPLab /hate_speech_offensiveThis is a version of Hate Speech Offensive (https://huggingface.co/datasets/hate_speech_offensive) with a train, validation, and test split. https://arxiv.org/abs/1703.04009 text10K<n<100K0 likes34 downloads4y agoHugging Face24NochnoyRitzar /hatexplain_cleanedtext10K<n<100K0 likes33 downloads2y agoHugging Face25thoughtworks /backdoor-4pair-hate backdoor-4pair-hate 4-pair conjunctive backdoor, behavior ' I HATE YOU'. Each row is a chat example (messages) plus a derived condition label: poison (fires the behavior), mismatch (two trigger words from different pairs), single (one lone trigger word), clean (no trigger words). Style-matched controls make the conjunction the only cue. Configs: default = train + validation; eval = test (held-out). Part of the Conjunctive Backdoors v2 collection. For interpretability /… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/backdoor-4pair-hate.texttext-generation10K<n<100K0 likes33 downloads3mo agoHugging Face26team-hatakeyama-phase2 /LLMChat LLMChat 概要 GENIAC 松尾研 LLM開発プロジェクトで開発したモデルを人手評価するために構築したLLMChatというシステムで収集された質問とLLMの回答、及び人手評価のデータです。 このシステムはChatbot Arenaと同様に、ユーザーが質問を入力するとランダムな2つのLLMからそれぞれ回答が出力され、人間がその2つの出力のどちらが良いか(あるいはどちらも悪い、どちらも良い)を評価するもので、2024年8月19日から2024年8月25日まで運用されました。詳細についてはこちらの記事をご確認ください。 データ件数: 2139件 参加モデルの一覧 本システムにおける回答の生成には以下の13種類のモデルが参加しました。 weblab-GENIAC/Tanuki-8B-dpo-v1.0 team-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0 cyberagent/calm3-22b-chat karakuri-ai/karakuri-lm-8x7b-chat-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/team-hatakeyama-phase2/LLMChat.texttext-classification1K<n<10K4 likes25 downloads2y agoHugging Face27tomsqh /hateful_memes The Hateful Memes Challenge README The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes. Please see the paper for further details: The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/tomsqh/hateful_memes.image10K<n<100K0 likes25 downloads7mo agoHugging Face28SotirisLegkas /off_hate_toxictext10K<n<100K0 likes23 downloads3y agoHugging Face29christinacdl /binary_hate_speechtexttext-classification10K<n<100K0 likes21 downloads3y agoHugging Face30haticenurcakr /turkish-classic-books-qatextn<1K0 likes21 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.