datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
0IEpEbL_Z3IlugLAO1sr_LLgoZ_nAJrCcrema-d-irony
CREMA-D-Irony
A controllable irony extension of CREMA-D,
introduced in SynIB: An Informational Bottleneck for Maximizing Synergy in Multimodal Learning
(arXiv:2606.09853 · code: github.com/kkontras/SynIB).
For a fraction α of clips, the original video is kept but its audio is replaced with a donor
clip carrying a contradicting emotion, and the sample is relabeled SAR (ironic). The label then
becomes recoverable only by combining audio and video — a controllable amount of… See the full description on the dataset page: https://huggingface.co/datasets/kkontras/crema-d-irony.task386_semeval_2018_task3_irony_detection
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task386_semeval_2018_task3_irony_detection
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task386_semeval_2018_task3_irony_detection.IronyTRHomepage: https://github.com/teghub/IronyTR
Labels:
0: non-ironic
1: ironic
EPIC_Irony
EPIC_Irony
paper: EPIC: Multi-Perspective Annotation of a Corpus of Irony at ACL 2023
Key features:
EPIC (English Perspectivist Irony Corpus) is an annotated corpus for irony analysis based on data perspectivism principles.
The corpus contains social media conversations in five regional varieties of English, annotated by contributors from corresponding countries.
The dataset explores the perspectives of annotators, taking into account their origin, age, and gender.… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/EPIC_Irony.multilingual_ironytask387_semeval_2018_task3_irony_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task387_semeval_2018_task3_irony_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task387_semeval_2018_task3_irony_classification.ptbr-irony-idioms-regionalism
Sotaques Digitais — Benchmark LLM Português Brasileiro
Dataset Summary
Sotaques Digitais é um benchmark de avaliação de competência pragmática e cultural para LLMs em português brasileiro. O dataset contém 90 cenários de teste distribuídos em três categorias linguísticas, construídos a partir de contextos reais (redes sociais, WhatsApp, atendimento ao cliente, avaliações de produto, ambiente de trabalho).
O benchmark foi desenvolvido para a pesquisa "Sotaques… See the full description on the dataset page: https://huggingface.co/datasets/ramondomiingos/ptbr-irony-idioms-regionalism.tweeteval-irony-mccirony-trtweet_eval-irony
Dataset Card for "tweet_eval-irony"
More Information needed
Note: This dataset was utilized for the evaluation of probability-based prompt selection techniques in the paper 'Improving Probability-based Prompt Selection Through Unified Evaluation and Analysis'. It differs from the actual benchmark dataset.
irony-trpp_distilbert_ft_tweet_ironyirony_tweetflan_combined_task387_semeval_2018_task3_irony_classificationirony-app-datasetsIronyEntityDatasetcustom_irony_test_set
