CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01darksyntax0 /conversational-sarcasm-benchmark Conversational Sarcasm Benchmark — Audio-Grounded, Metadata-Only A benchmark of 1,168 conversational sarcasm units drawn from 64 English-language YouTube videos (predominantly stand-up comedy and comedic conversation). Every unit pairs a short target utterance with the preceding context that makes its figurative reading available, and carries a categorical label plus a free-text rationale. This repository contains no audio. It ships annotations, transcriptions, and the source… See the full description on the dataset page: https://huggingface.co/datasets/darksyntax0/conversational-sarcasm-benchmark.tabularaudio-classification1K<n<10K0 likes73 downloads28d agoHugging Face02mounikaiiith /Telugu-SarcasmDo cite the below references for using the dataset: @article{marreddy2022resource, title={Am I a Resource-Poor Language? Data Sets, Embeddings, Models and Analysis for four different NLP tasks in Telugu Language}, author={Marreddy, Mounika and Oota, Subba Reddy and Vakada, Lakshmi Sireesha and Chinni, Venkata Charan and Mamidi, Radhika}, journal={Transactions on Asian and Low-Resource Language Information Processing}, publisher={ACM New York, NY} } @article{marreddy2022multi… See the full description on the dataset page: https://huggingface.co/datasets/mounikaiiith/Telugu-Sarcasm.text10K<n<100K2 likes41 downloads4y agoHugging Face03daniel2588 /sarcasm Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/daniel2588/sarcasm.tabular1M<n<10M2 likes35 downloads3y agoHugging Face04Dalilame /Multi-Sarcasmgated Multi-Sarcasm Dataset Multi-Sarcasm is a curated conversational dataset containing 23,715 parent–reply pairs, combining Reddit and chatbot dialogues. Each instance includes a comment (comment), its parent comment (parent_comment), a binary label (label) for sarcasm, and a graded sarcasm level (sarcasm_level). The dataset was re-annotated using a hybrid LLM-assisted pre-annotation followed by systematic human validation, ensuring consistency, diversity, and reliability.… See the full description on the dataset page: https://huggingface.co/datasets/Dalilame/Multi-Sarcasm.tabulartext-classification10K<n<100K5 likes33 downloads6mo agoHugging Face05helinivan /sarcasm_headlines_multilingual Dataset Card for Multilingual Sarcasm Detection Dataset Summary Dataset consists of news article headlines in Dutch, English and Italian. The news article headlines are both from actual news sources and sarcastic/satirical newspapers. The news article is determined sarcastic/non-sarcastic based on the news article source. The sources of news articles are: The Huffington Post (en, non-sarcastic) The Onion (en, sarcastic) NOS (nl, non-sarcastic) De Speld (nl, sarcastic) Il… See the full description on the dataset page: https://huggingface.co/datasets/helinivan/sarcasm_headlines_multilingual.tabular10K<n<100K1 likes30 downloads4y agoHugging Face06nikesh66 /Sarcasm-dataset Sarcasm Dataset This dataset contains sarcastic sentence along with their binary label Dataset Desciption: Number of Rows: 5,000 Number of Columns: 2 Column Names: 'Tweet', 'Slang (yes/no)' Description: The dataset contains tweets annotated for the use of slang. It includes a binary label ('yes' or 'no') indicating the presence of slang in each tweet. text1K<n<10K4 likes28 downloads3y agoHugging Face07processvenue /SARCASM_VS_NON_SARCASMtexttext-classification1K<n<10K0 likes27 downloads10mo agoHugging Face08bushra242khan /Sarcasm-dataset Sarcasm Dataset This dataset contains sarcastic sentence along with their binary label Dataset Desciption: Number of Rows: 5,000 Number of Columns: 2 Column Names: 'Tweet', 'Slang (yes/no)' Description: The dataset contains tweets annotated for the use of slang. It includes a binary label ('yes' or 'no') indicating the presence of slang in each tweet. text1K<n<10K0 likes25 downloads2d agoHugging Face09patrickjamesmarcellana /authentic-filipino-sarcasm-detection Authentic Filipino Sarcasm Detection Dataset This dataset is composed of Filipino sarcastic and non-sarcastic tweets scraped from X (formerly Twitter), divided into two categories: politics and entertainment. Dataset Size The dataset is composed of 1,000 tweets, 500 for each domain of politics and entertainment. Rows Each row is an instance of a tweet, constrained with X's limitation of 280 characters. Columns text: the tweet content label:… See the full description on the dataset page: https://huggingface.co/datasets/patrickjamesmarcellana/authentic-filipino-sarcasm-detection.texttext-classification1K<n<10K0 likes22 downloads1y agoHugging Face10rishikesh /mini-sarcasm-datatext1K<n<10K2 likes16 downloads3y agoHugging Face11patrickjamesmarcellana /synthetic-filipino-sarcasm-detection Synthetic and Limited Real-World Filipino Sarcasm Detection Dataset This dataset is composed of Filipino sarcastic and non-sarcastic tweets, divided into two categories: LLM-generated (synthetic) data and real-world data. Synthetic Data Information Two large language models were used to generate sarcastic and non-sarcastic tweets for the dataset: GPT-4o and Gemini 2.0 Flash. The dataset is composed of 504 sarcastic tweets (252 for each LLM) and 504 non-sarcastic tweets… See the full description on the dataset page: https://huggingface.co/datasets/patrickjamesmarcellana/synthetic-filipino-sarcasm-detection.texttext-classification1K<n<10K0 likes14 downloads1y agoHugging Face12siddhant4583agarwal /sarcasm-detection-datasettext1K<n<10K0 likes11 downloads2y agoHugging Face13DrDavis /sarcasm-englishtext1K<n<10K1 likes8 downloads11mo agoHugging Face14anonymous813ker /sarcasm-synthetictexttext-classification1K<n<10K0 likes8 downloads4mo agoHugging Face15daniel2588 /sarcasmdatatext10K<n<100K0 likes7 downloads3y agoHugging Face16maliha /sarcasm-explain-5kgated SarcasmExplain-5K Dataset Description SarcasmExplain-5K is a balanced dataset of 5,000 Reddit sarcasm instances annotated with five complementary natural language explanation types, generated via GPT-4 and validated through human evaluation. Created by: Maliha Binte Mamun Year: 2025 License: CC BY 4.0 GitHub: https://github.com/maliha-usui/sarcasm-explain-5k 🔒 Access Complete 3 annotation forms to download: 👉… See the full description on the dataset page: https://huggingface.co/datasets/maliha/sarcasm-explain-5k.texttext-classification1K<n<10K0 likes7 downloads7mo agoHugging Face17salsabilahasna /Sarcasm_Dataset NLP DATASET-Sarcasm Source code : https://www.kaggle.com/datasets/toygarr/datasets-for-natural-language-processing text10K<n<100K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.