datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nlu_evaluation_dataRaw part of NLU Evaluation Data. It contains 25 715 non-empty examples (original dataset has 25716 examples) from 68 unique intents belonging to 18 scenarios.NLU-Sentiment-Analysis
SEA Sentiment Analysis
SEA Sentiment Analysis evaluates a model's ability to identify the sentiment polarity of a text. It is sampled from NusaX for Indonesian, Javanese, and Sundanese, IndicSentiment for Tamil, Wisesight Sentiment for Thai, and UIT-VSFC for Vietnamese.
Supported Tasks and Leaderboards
SEA Sentiment Analysis is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Sentiment-Analysis.NLU-Question-Answering
SEA Question Answering
SEA Question Answering evaluates a model's ability to predict a contiguous span of characters that answers the question about a given passage. It is sampled from TyDi QA-GoldP for Indonesian, IndicQA for Tamil, and XQuaD for Thai and Vietnamese.
Supported Tasks and Leaderboards
SEA Question Answering is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Question-Answering.NLU-Belebele-MCQAsnips_built_in_intents
Dataset Card for Snips Built In Intents
Dataset Summary
Snips' built in intents dataset was initially used to compare different voice assistants and released as a public dataset hosted at
https://github.com/sonos/nlu-benchmark in folder 2016-12-built-in-intents. The dataset contains 328 utterances over 10 intent classes.
A related Medium post is https://medium.com/snips-ai/benchmarking-natural-language-understanding-systems-d35be6ce568d.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sonos-nlu-benchmark/snips_built_in_intents.NLU-Metaphor
SEA Metaphor
SEA Metaphor evaluates a model's ability to interpret paired figurative phrases with divergent meanings. It is sampled from Multilingual-Fig-QA for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Metaphor is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Javanese (jv)
Sundanese (su)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Metaphor.multi3-nlu
Dataset Card for Multi3NLU++
Dataset Summary
Please access the dataset using
git clone https://huggingface.co/datasets/uoe-nlp/multi3-nlu/
Multi3NLU++ consists of 3080 utterances per language representing challenges in building multilingual multi-intent multi-domain task-oriented dialogue systems. The domains include banking and hotels. There are 62 unique intents.
Supported Tasks and Leaderboards
multi-label intent detection
slot filling
cross-lingual… See the full description on the dataset page: https://huggingface.co/datasets/uoe-nlp/multi3-nlu.FDA_for_NLUMultiPICo
Dataset Summary
MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on).
Supported Tasks and Leaderboards
Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.kor_nlu The dataset contains data for bechmarking korean models on NLI and STSdetails_NLUHOPOE__test-case-2
Dataset Card for Evaluation run of NLUHOPOE/test-case-2
Dataset automatically created during the evaluation run of model NLUHOPOE/test-case-2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NLUHOPOE__test-case-2.details_NLUHOPOE__test-case-0
Dataset Card for Evaluation run of NLUHOPOE/test-case-0
Dataset automatically created during the evaluation run of model NLUHOPOE/test-case-0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NLUHOPOE__test-case-0.EPIC
Dataset Card for EPICorpus
Dataset Summary
EPIC (English Perspectivist Irony Corpus) is a disaggregated English corpus for irony detection, containing 3,000 pairs of short conversations (posts-replies) from Twitter and Reddit, along with the demographic information of each annotator (age, nationality, gender, and so on).
Supported Tasks and Leaderboards
Irony classification task using soft labels (i.e., distribution of annotations) or hard labels… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/EPIC.citizen_nlu
Dataset Card for citizen_nlu
Dataset Summary
NeuralSpace strives to provide AutoNLP text and speech services, especially for low-resource languages. One of the major services provided by NeuralSpace on its platform is the “Language Understanding” service, where you can build, train and deploy your NLU model to recognize intents and entities with minimal code and just a few clicks.
The initiative of this challenge is created with the purpose of sparkling AI applications to… See the full description on the dataset page: https://huggingface.co/datasets/neuralspace/citizen_nlu.details_NLUHOPOE__test-case-1
Dataset Card for Evaluation run of NLUHOPOE/test-case-1
Dataset automatically created during the evaluation run of model NLUHOPOE/test-case-1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NLUHOPOE__test-case-1.details_NLUHOPOE__experiment2-cause
Dataset Card for Evaluation run of NLUHOPOE/experiment2-cause
Dataset automatically created during the evaluation run of model NLUHOPOE/experiment2-cause on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NLUHOPOE__experiment2-cause.details_NLUHOPOE__experiment2-cause-non
Dataset Card for Evaluation run of NLUHOPOE/experiment2-cause-non
Dataset automatically created during the evaluation run of model NLUHOPOE/experiment2-cause-non on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NLUHOPOE__experiment2-cause-non.persian-nlu
Persian NLU
Dataset Summary
The Persian NLU Benchmark is a curated collection of existing Persian datasets designed to evaluate Natural Language Understanding (NLU) capabilities across a diverse range of tasks. It provides a unified benchmark suite to assess different cognitive aspects of large language models (LLMs) in Persian.
This benchmark includes the following tasks and datasets:
Text Classification:
Synthetic Persian Tone
SID
Natural Language Inference… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/persian-nlu.details_NLUHOPOE__experiment2-cause-non-qLoRa
Dataset Card for Evaluation run of NLUHOPOE/experiment2-cause-non-qLoRa
Dataset automatically created during the evaluation run of model NLUHOPOE/experiment2-cause-non-qLoRa on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NLUHOPOE__experiment2-cause-non-qLoRa.sentinelng-data-nlu
SentinelNG NLU Dataset
This dataset repository contains SentinelNG natural-language understanding resources for crop, health, security, and other intents. The tree includes intent text files, conversational JSON/JSONL resources, and an nigerian_nlu.ftz model artifact.
Contents
The repository includes intent-oriented text files such as crop_intent.txt, health_intent.txt, security_intent.txt, and other_intent.txt, together with conversational resources and… See the full description on the dataset page: https://huggingface.co/datasets/MR-CODESPIKE/sentinelng-data-nlu.details_NLUHOPOE__Mistral-7B-length-100000
Dataset Card for Evaluation run of NLUHOPOE/Mistral-7B-length-100000
Dataset automatically created during the evaluation run of model NLUHOPOE/Mistral-7B-length-100000 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NLUHOPOE__Mistral-7B-length-100000.formosa-nlu-synth-v1
FormosaNLU Synth
FormosaNLU Synth 是一份以正體中文(台灣,zh-TW)為主的口語 NLU
synthetic training dataset,涵蓋 60 種 intent 與 55 種 slot type。資料由
本機 open-weight teacher 產生,經 deterministic F1–F6 filters 與不同家族
independent judge(F7)稽核後,發布 3,754 筆 training rows。
本資料集對應的完整程式碼、決策紀錄與實驗報告:
kuotunyu/FormosaNLU-Synth。
內容
data/train.jsonl 3,754 rows
schema.json JSON Schema
release_manifest.json 來源 artifact、SHA-256、筆數與版本
每筆資料包含:
欄位
說明
id
穩定 synthetic sample ID
utt… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/formosa-nlu-synth-v1.NLU-Evaluation-Data-en-de
NLU Evaluation Data - English and German
A labeled English and German language multi-domain dataset (21 domains) with 25K user utterances for human-robot interaction.
This dataset is collected and annotated for evaluating NLU services and platforms.
The detailed paper on this dataset can be found at arXiv.org:
Benchmarking Natural Language Understanding Services for building Conversational Agents
The dataset builds on the annotated data of the xliuhw/NLU-Evaluation-Data
repository.… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-Evaluation-Data-en-de.details_NLUHOPOE__test-case-3
Dataset Card for Evaluation run of NLUHOPOE/test-case-3
Dataset automatically created during the evaluation run of model NLUHOPOE/test-case-3 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NLUHOPOE__test-case-3.egyptian-nlu
Egyptian Arabic voice-assistant NLU — dataset
Egyptian Arabic commands paired with intent + slot annotations, in the schema of
Amazon MASSIVE (60 intents, 55 slot types).
Two parts, and the difference matters:
File
Rows
Origin
egy_test.jsonl
200
Written and annotated by hand by a native Egyptian speaker — the benchmark
egy_synth_train.jsonl
6,509
LLM-generated Egyptian rewrites of MASSIVE ar-SA training items
egy_synth_dev.jsonl
730
Same, held out by seed… See the full description on the dataset page: https://huggingface.co/datasets/Alhasan/egyptian-nlu.details_NLUHOPOE__test-case-5
Dataset Card for Evaluation run of NLUHOPOE/test-case-5
Dataset automatically created during the evaluation run of model NLUHOPOE/test-case-5 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NLUHOPOE__test-case-5.details_NLUHOPOE__Mistral-7B-attention-100000
Dataset Card for Evaluation run of NLUHOPOE/Mistral-7B-attention-100000
Dataset automatically created during the evaluation run of model NLUHOPOE/Mistral-7B-attention-100000 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NLUHOPOE__Mistral-7B-attention-100000.details_NLUHOPOE__test-case-6
Dataset Card for Evaluation run of NLUHOPOE/test-case-6
Dataset automatically created during the evaluation run of model NLUHOPOE/test-case-6 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NLUHOPOE__test-case-6.incar-nlu-judge-out-v31aveni-bench-polyai-nlu
AveniBench: PolyAI NLU++
PolyAI NLU++ split used in the AveniBench.
License
This dataset is made available under the CC-BY-4.0 license.
Citation
AveniBench
TDB
PolyAI NLU++
@inproceedings{casanueva-etal-2022-nlu,
title = "{NLU}++: A Multi-Label, Slot-Rich, Generalisable Dataset for Natural Language Understanding in Task-Oriented Dialogue",
author = "Casanueva, Inigo and
Vuli{\'c}, Ivan and
Spithourakis, Georgios and
Budzianowski… See the full description on the dataset page: https://huggingface.co/datasets/aveni-ai/aveni-bench-polyai-nlu.
