datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiPICo
Dataset Summary
MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on).
Supported Tasks and Leaderboards
Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.EPIC
Dataset Card for EPICorpus
Dataset Summary
EPIC (English Perspectivist Irony Corpus) is a disaggregated English corpus for irony detection, containing 3,000 pairs of short conversations (posts-replies) from Twitter and Reddit, along with the demographic information of each annotator (age, nationality, gender, and so on).
Supported Tasks and Leaderboards
Irony classification task using soft labels (i.e., distribution of annotations) or hard labels… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/EPIC.NLU-Evaluation-Data-en-de
NLU Evaluation Data - English and German
A labeled English and German language multi-domain dataset (21 domains) with 25K user utterances for human-robot interaction.
This dataset is collected and annotated for evaluating NLU services and platforms.
The detailed paper on this dataset can be found at arXiv.org:
Benchmarking Natural Language Understanding Services for building Conversational Agents
The dataset builds on the annotated data of the xliuhw/NLU-Evaluation-Data
repository.… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-Evaluation-Data-en-de.NLU-few-shot-benchmark-en-de
NLU Few-shot Benchmark - English and German
This is a few-shot training dataset from the domain of human-robot interaction.
It contains texts in German and English language with 64 different utterances (classes).
Each utterance (class) has exactly 20 samples in the training set.
This leads to a total of 1280 different training samples.
The dataset is intended to benchmark the intent classifiers of chat bots in English and especially in German language.
We are building on our… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-few-shot-benchmark-en-de.DIARC-embodied-nlu-styled-4k
DIARC-LLM-Parser-Embodied-NLU-Styled-4K
This dataset contains about ~4k utterances together with their semantic parses as interpretable by the DIARC cognitive robotic architecture.
The parses are meant to capture the speech-theoretic aspects of NL and parse the intent, referents, and descriptors in the utterance.
This dataset is one in a set of datasets. For this particular one, we programmatically built 127 utterances and semantics that are groundable in a robotic architecture… See the full description on the dataset page: https://huggingface.co/datasets/vsarathy/DIARC-embodied-nlu-styled-4k.kor_nlu_hufscitizen_nluDIARC-embodied-nlu-styled-4k-with-contextsheng_nlu
Common User Intentions
Greetings
Wasemaje
uko aje btw
oyah...
Form
Alafu niaje
Poa Sana Mambo
Niko poa
Pia Mimi Niko salama
Hope siku yako iko poa
Siko poa kabisa
Nimekuwa poa
Umeshindaje
Hope uko poa
uko poa
Sasa
Vipi vipi
Niko salama
..its been long.
Nko fiti
niko fiti
Nmeamka fity..
Vipi
Unasemaje
Aaaah...itakuaje sasaa..
.iz vipi..itakuaje..
Form ni gani bro...
iz vipi
Affirm
Hapo sawa...
Fty
sai
Hio si ni better hadi
Imebidi.
Eeeh mazee
mazeee
Fity… See the full description on the dataset page: https://huggingface.co/datasets/JeunesseAfricaine/sheng_nlu.NLU-Courseworkfunction-calling-small-openai
