datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NLU-Sentiment-Analysis
SEA Sentiment Analysis
SEA Sentiment Analysis evaluates a model's ability to identify the sentiment polarity of a text. It is sampled from NusaX for Indonesian, Javanese, and Sundanese, IndicSentiment for Tamil, Wisesight Sentiment for Thai, and UIT-VSFC for Vietnamese.
Supported Tasks and Leaderboards
SEA Sentiment Analysis is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Sentiment-Analysis.NLU-Question-Answering
SEA Question Answering
SEA Question Answering evaluates a model's ability to predict a contiguous span of characters that answers the question about a given passage. It is sampled from TyDi QA-GoldP for Indonesian, IndicQA for Tamil, and XQuaD for Thai and Vietnamese.
Supported Tasks and Leaderboards
SEA Question Answering is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Question-Answering.NLU-Belebele-MCQAsnips_built_in_intents
Dataset Card for Snips Built In Intents
Dataset Summary
Snips' built in intents dataset was initially used to compare different voice assistants and released as a public dataset hosted at
https://github.com/sonos/nlu-benchmark in folder 2016-12-built-in-intents. The dataset contains 328 utterances over 10 intent classes.
A related Medium post is https://medium.com/snips-ai/benchmarking-natural-language-understanding-systems-d35be6ce568d.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sonos-nlu-benchmark/snips_built_in_intents.NLU-Metaphor
SEA Metaphor
SEA Metaphor evaluates a model's ability to interpret paired figurative phrases with divergent meanings. It is sampled from Multilingual-Fig-QA for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Metaphor is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Javanese (jv)
Sundanese (su)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Metaphor.MultiPICo
Dataset Summary
MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on).
Supported Tasks and Leaderboards
Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.EPIC
Dataset Card for EPICorpus
Dataset Summary
EPIC (English Perspectivist Irony Corpus) is a disaggregated English corpus for irony detection, containing 3,000 pairs of short conversations (posts-replies) from Twitter and Reddit, along with the demographic information of each annotator (age, nationality, gender, and so on).
Supported Tasks and Leaderboards
Irony classification task using soft labels (i.e., distribution of annotations) or hard labels… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/EPIC.egyptian-nlu
Egyptian Arabic voice-assistant NLU — dataset
Egyptian Arabic commands paired with intent + slot annotations, in the schema of
Amazon MASSIVE (60 intents, 55 slot types).
Two parts, and the difference matters:
File
Rows
Origin
egy_test.jsonl
200
Written and annotated by hand by a native Egyptian speaker — the benchmark
egy_synth_train.jsonl
6,509
LLM-generated Egyptian rewrites of MASSIVE ar-SA training items
egy_synth_dev.jsonl
730
Same, held out by seed… See the full description on the dataset page: https://huggingface.co/datasets/Alhasan/egyptian-nlu.formosa-nlu-synth-v1
FormosaNLU Synth
FormosaNLU Synth 是一份以正體中文(台灣,zh-TW)為主的口語 NLU
synthetic training dataset,涵蓋 60 種 intent 與 55 種 slot type。資料由
本機 open-weight teacher 產生,經 deterministic F1–F6 filters 與不同家族
independent judge(F7)稽核後,發布 3,754 筆 training rows。
本資料集對應的完整程式碼、決策紀錄與實驗報告:
kuotunyu/FormosaNLU-Synth。
內容
data/train.jsonl 3,754 rows
schema.json JSON Schema
release_manifest.json 來源 artifact、SHA-256、筆數與版本
每筆資料包含:
欄位
說明
id
穩定 synthetic sample ID
utt… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/formosa-nlu-synth-v1.sentinelng-data-nlu
SentinelNG NLU Dataset
This dataset repository contains SentinelNG natural-language understanding resources for crop, health, security, and other intents. The tree includes intent text files, conversational JSON/JSONL resources, and an nigerian_nlu.ftz model artifact.
Contents
The repository includes intent-oriented text files such as crop_intent.txt, health_intent.txt, security_intent.txt, and other_intent.txt, together with conversational resources and… See the full description on the dataset page: https://huggingface.co/datasets/MR-CODESPIKE/sentinelng-data-nlu.NLU-Evaluation-Data-en-de
NLU Evaluation Data - English and German
A labeled English and German language multi-domain dataset (21 domains) with 25K user utterances for human-robot interaction.
This dataset is collected and annotated for evaluating NLU services and platforms.
The detailed paper on this dataset can be found at arXiv.org:
Benchmarking Natural Language Understanding Services for building Conversational Agents
The dataset builds on the annotated data of the xliuhw/NLU-Evaluation-Data
repository.… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-Evaluation-Data-en-de.aveni-bench-polyai-nlu
AveniBench: PolyAI NLU++
PolyAI NLU++ split used in the AveniBench.
License
This dataset is made available under the CC-BY-4.0 license.
Citation
AveniBench
TDB
PolyAI NLU++
@inproceedings{casanueva-etal-2022-nlu,
title = "{NLU}++: A Multi-Label, Slot-Rich, Generalisable Dataset for Natural Language Understanding in Task-Oriented Dialogue",
author = "Casanueva, Inigo and
Vuli{\'c}, Ivan and
Spithourakis, Georgios and
Budzianowski… See the full description on the dataset page: https://huggingface.co/datasets/aveni-ai/aveni-bench-polyai-nlu.Palette_nlu_service
Persian Sales NLU (seed)
Synthetic + templated Persian sales dialogues for NLU training.
Splits: train.jsonl, val.jsonl, test.jsonl
Each line: text, tokens, intent, slots (BIO), keep_mask, search_query.
incar-nlu-judge-out-v31greek-nlu-bench
Greek NLU Benchmark
Frozen (v1.0) Greek language-understanding benchmark — 9,751 items / 18 tasks, one split per task. Native-verified and train/eval-disjoint. Built to detect CPT/SFT gains and regressions on Greek, weighted toward morphology-sensitive tasks.
Item schema
One JSON object per line; identical envelope across tasks:
{
"id": "el-mcq-000042", "task": "mcq", "language": "el", "version": "1.0",
"tags": {"category": "linguistic", "phenomenon":… See the full description on the dataset page: https://huggingface.co/datasets/KIEFERSA/greek-nlu-bench.NLU-few-shot-benchmark-en-de
NLU Few-shot Benchmark - English and German
This is a few-shot training dataset from the domain of human-robot interaction.
It contains texts in German and English language with 64 different utterances (classes).
Each utterance (class) has exactly 20 samples in the training set.
This leads to a total of 1280 different training samples.
The dataset is intended to benchmark the intent classifiers of chat bots in English and especially in German language.
We are building on our… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-few-shot-benchmark-en-de.Palette_nlu_service
Persian Sales NLU (seed)
Synthetic + templated Persian sales dialogues for NLU training.
Splits: train.jsonl, val.jsonl, test.jsonl
Each line: text, tokens, intent, slots (BIO), keep_mask, search_query.
indo-nlu-entailmen
Dataset Card for "indo-nlu-entailmen"
More Information needed
Chess_piecesnew_nlu_tts3
Dataset Card for "new_nlu_tts3"
More Information needed
dataset_VMLU_for_bloomparsi-nluNLU-Redact-PII-v1
Synthetic Dataset Data Card
This document provides an overview of the synthetic dataset generated for testing redaction and anonymization pipelines. It outlines the data generation process, the variety in data formats, ethical considerations, and the impact of complex invalid formats on model quality.
Overview
The synthetic dataset is created using a suite of generators that produce both valid and intentionally invalid formats for sensitive data such as names, card… See the full description on the dataset page: https://huggingface.co/datasets/darkmatter2222/NLU-Redact-PII-v1.dataset_dhnl_qna_v2audiologynlu-covidFrench benchmark of NLU services for employee support use case during covid-19 pandemic.
These datasets were created by the Wikit team in order to compare the performances of NLU tools on the French language.
The dataset use case is employee support during the covid 19 pandemic. The intents were defined to answer department employees' questions on the evolution of work conditions related to the crisis.
The training_dataset.csv file contains training utterances with associated intent used to… See the full description on the dataset page: https://huggingface.co/datasets/Wikit/nlu-covid.ru_nlu_intentDIARC-embodied-nlu-styled-4k
DIARC-LLM-Parser-Embodied-NLU-Styled-4K
This dataset contains about ~4k utterances together with their semantic parses as interpretable by the DIARC cognitive robotic architecture.
The parses are meant to capture the speech-theoretic aspects of NL and parse the intent, referents, and descriptors in the utterance.
This dataset is one in a set of datasets. For this particular one, we programmatically built 127 utterances and semantics that are groundable in a robotic architecture… See the full description on the dataset page: https://huggingface.co/datasets/vsarathy/DIARC-embodied-nlu-styled-4k.kor_nlu_hufsnlu_stsv_test
