datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
egyptian-nlu
Egyptian Arabic voice-assistant NLU — dataset
Egyptian Arabic commands paired with intent + slot annotations, in the schema of
Amazon MASSIVE (60 intents, 55 slot types).
Two parts, and the difference matters:
File
Rows
Origin
egy_test.jsonl
200
Written and annotated by hand by a native Egyptian speaker — the benchmark
egy_synth_train.jsonl
6,509
LLM-generated Egyptian rewrites of MASSIVE ar-SA training items
egy_synth_dev.jsonl
730
Same, held out by seed… See the full description on the dataset page: https://huggingface.co/datasets/Alhasan/egyptian-nlu.formosa-nlu-synth-v1
FormosaNLU Synth
FormosaNLU Synth 是一份以正體中文(台灣,zh-TW)為主的口語 NLU
synthetic training dataset,涵蓋 60 種 intent 與 55 種 slot type。資料由
本機 open-weight teacher 產生,經 deterministic F1–F6 filters 與不同家族
independent judge(F7)稽核後,發布 3,754 筆 training rows。
本資料集對應的完整程式碼、決策紀錄與實驗報告:
kuotunyu/FormosaNLU-Synth。
內容
data/train.jsonl 3,754 rows
schema.json JSON Schema
release_manifest.json 來源 artifact、SHA-256、筆數與版本
每筆資料包含:
欄位
說明
id
穩定 synthetic sample ID
utt… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/formosa-nlu-synth-v1.aveni-bench-polyai-nlu
AveniBench: PolyAI NLU++
PolyAI NLU++ split used in the AveniBench.
License
This dataset is made available under the CC-BY-4.0 license.
Citation
AveniBench
TDB
PolyAI NLU++
@inproceedings{casanueva-etal-2022-nlu,
title = "{NLU}++: A Multi-Label, Slot-Rich, Generalisable Dataset for Natural Language Understanding in Task-Oriented Dialogue",
author = "Casanueva, Inigo and
Vuli{\'c}, Ivan and
Spithourakis, Georgios and
Budzianowski… See the full description on the dataset page: https://huggingface.co/datasets/aveni-ai/aveni-bench-polyai-nlu.Palette_nlu_service
Persian Sales NLU (seed)
Synthetic + templated Persian sales dialogues for NLU training.
Splits: train.jsonl, val.jsonl, test.jsonl
Each line: text, tokens, intent, slots (BIO), keep_mask, search_query.
greek-nlu-bench
Greek NLU Benchmark
Frozen (v1.0) Greek language-understanding benchmark — 9,751 items / 18 tasks, one split per task. Native-verified and train/eval-disjoint. Built to detect CPT/SFT gains and regressions on Greek, weighted toward morphology-sensitive tasks.
Item schema
One JSON object per line; identical envelope across tasks:
{
"id": "el-mcq-000042", "task": "mcq", "language": "el", "version": "1.0",
"tags": {"category": "linguistic", "phenomenon":… See the full description on the dataset page: https://huggingface.co/datasets/KIEFERSA/greek-nlu-bench.Palette_nlu_service
Persian Sales NLU (seed)
Synthetic + templated Persian sales dialogues for NLU training.
Splits: train.jsonl, val.jsonl, test.jsonl
Each line: text, tokens, intent, slots (BIO), keep_mask, search_query.
dataset_VMLU_for_bloomNLU-Redact-PII-v1
Synthetic Dataset Data Card
This document provides an overview of the synthetic dataset generated for testing redaction and anonymization pipelines. It outlines the data generation process, the variety in data formats, ethical considerations, and the impact of complex invalid formats on model quality.
Overview
The synthetic dataset is created using a suite of generators that produce both valid and intentionally invalid formats for sensitive data such as names, card… See the full description on the dataset page: https://huggingface.co/datasets/darkmatter2222/NLU-Redact-PII-v1.dataset_dhnl_qna_v2nlu-covidFrench benchmark of NLU services for employee support use case during covid-19 pandemic.
These datasets were created by the Wikit team in order to compare the performances of NLU tools on the French language.
The dataset use case is employee support during the covid 19 pandemic. The intents were defined to answer department employees' questions on the evolution of work conditions related to the crisis.
The training_dataset.csv file contains training utterances with associated intent used to… See the full description on the dataset page: https://huggingface.co/datasets/Wikit/nlu-covid.nlu_stsv_testdataset_dhnl_qnanlu_qna_demoshiji-70liezhuannlu_qna_demo2vlmu_eval_935rowsZaloAI_QnAqchv_2021_qnaVLMU_PhoGPTdataset_NLU_QnAnlu_promptsVLMU_PhoGPT_2nlu_stsv_test2qchv_2021_qna_v2qna_qchv_2021testq_avlmu_train_eval_100rowsqchv_2021test_format
