datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
authentic-filipino-sarcasm-detection
Authentic Filipino Sarcasm Detection Dataset
This dataset is composed of Filipino sarcastic and non-sarcastic tweets scraped from X (formerly Twitter), divided into two categories: politics and entertainment.
Dataset Size
The dataset is composed of 1,000 tweets, 500 for each domain of politics and entertainment.
Rows
Each row is an instance of a tweet, constrained with X's limitation of 280 characters.
Columns
text:
the tweet content
label:… See the full description on the dataset page: https://huggingface.co/datasets/patrickjamesmarcellana/authentic-filipino-sarcasm-detection.synthetic-filipino-sarcasm-detection
Synthetic and Limited Real-World Filipino Sarcasm Detection Dataset
This dataset is composed of Filipino sarcastic and non-sarcastic tweets, divided into two categories: LLM-generated (synthetic) data and real-world data.
Synthetic Data Information
Two large language models were used to generate sarcastic and non-sarcastic tweets for the dataset: GPT-4o and Gemini 2.0 Flash. The dataset is composed of 504 sarcastic tweets (252 for each LLM) and 504 non-sarcastic tweets… See the full description on the dataset page: https://huggingface.co/datasets/patrickjamesmarcellana/synthetic-filipino-sarcasm-detection.sarcasm-detection-dataset
