CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aisingapore /NLU-Sentiment-Analysisgated SEA Sentiment Analysis SEA Sentiment Analysis evaluates a model's ability to identify the sentiment polarity of a text. It is sampled from NusaX for Indonesian, Javanese, and Sundanese, IndicSentiment for Tamil, Wisesight Sentiment for Thai, and UIT-VSFC for Vietnamese. Supported Tasks and Leaderboards SEA Sentiment Analysis is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Sentiment-Analysis.texttext-generation1K<n<10K0 likes2.6k downloads9mo agoHugging Face02aisingapore /NLU-Question-Answeringgated SEA Question Answering SEA Question Answering evaluates a model's ability to predict a contiguous span of characters that answers the question about a given passage. It is sampled from TyDi QA-GoldP for Indonesian, IndicQA for Tamil, and XQuaD for Thai and Vietnamese. Supported Tasks and Leaderboards SEA Question Answering is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Question-Answering.texttext-generation1K<n<10K0 likes1.9k downloads9mo agoHugging Face03aisingapore /NLU-Belebele-MCQAgatedtext10K<n<100K0 likes1.6k downloads9mo agoHugging Face04sonos-nlu-benchmark /snips_built_in_intents Dataset Card for Snips Built In Intents Dataset Summary Snips' built in intents dataset was initially used to compare different voice assistants and released as a public dataset hosted at https://github.com/sonos/nlu-benchmark in folder 2016-12-built-in-intents. The dataset contains 328 utterances over 10 intent classes. A related Medium post is https://medium.com/snips-ai/benchmarking-natural-language-understanding-systems-d35be6ce568d. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sonos-nlu-benchmark/snips_built_in_intents.texttext-classificationn<1K14 likes1.1k downloads2y agoHugging Face05aisingapore /NLU-Metaphorgated SEA Metaphor SEA Metaphor evaluates a model's ability to interpret paired figurative phrases with divergent meanings. It is sampled from Multilingual-Fig-QA for Indonesian, Javanese, and Sundanese. Supported Tasks and Leaderboards SEA Metaphor is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore. Languages Indonesian (id) Javanese (jv) Sundanese (su) Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Metaphor.texttext-generation1K<n<10K0 likes921 downloads9mo agoHugging Face06Multilingual-Perspectivist-NLU /MultiPICo Dataset Summary MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.tabular10K<n<100K6 likes230 downloads2y agoHugging Face07Multilingual-Perspectivist-NLU /EPIC Dataset Card for EPICorpus Dataset Summary EPIC (English Perspectivist Irony Corpus) is a disaggregated English corpus for irony detection, containing 3,000 pairs of short conversations (posts-replies) from Twitter and Reddit, along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/EPIC.tabulartext-classification10K<n<100K2 likes161 downloads2y agoHugging Face08Alhasan /egyptian-nlu Egyptian Arabic voice-assistant NLU — dataset Egyptian Arabic commands paired with intent + slot annotations, in the schema of Amazon MASSIVE (60 intents, 55 slot types). Two parts, and the difference matters: File Rows Origin egy_test.jsonl 200 Written and annotated by hand by a native Egyptian speaker — the benchmark egy_synth_train.jsonl 6,509 LLM-generated Egyptian rewrites of MASSIVE ar-SA training items egy_synth_dev.jsonl 730 Same, held out by seed… See the full description on the dataset page: https://huggingface.co/datasets/Alhasan/egyptian-nlu.texttext-generation1K<n<10K0 likes60 downloads21h agoHugging Face09steven0226 /formosa-nlu-synth-v1 FormosaNLU Synth FormosaNLU Synth 是一份以正體中文(台灣,zh-TW)為主的口語 NLU synthetic training dataset,涵蓋 60 種 intent 與 55 種 slot type。資料由 本機 open-weight teacher 產生,經 deterministic F1–F6 filters 與不同家族 independent judge(F7)稽核後,發布 3,754 筆 training rows。 本資料集對應的完整程式碼、決策紀錄與實驗報告: kuotunyu/FormosaNLU-Synth。 內容 data/train.jsonl 3,754 rows schema.json JSON Schema release_manifest.json 來源 artifact、SHA-256、筆數與版本 每筆資料包含: 欄位 說明 id 穩定 synthetic sample ID utt… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/formosa-nlu-synth-v1.texttext-classification1K<n<10K0 likes49 downloads1mo agoHugging Face10MR-CODESPIKE /sentinelng-data-nlu SentinelNG NLU Dataset This dataset repository contains SentinelNG natural-language understanding resources for crop, health, security, and other intents. The tree includes intent text files, conversational JSON/JSONL resources, and an nigerian_nlu.ftz model artifact. Contents The repository includes intent-oriented text files such as crop_intent.txt, health_intent.txt, security_intent.txt, and other_intent.txt, together with conversational resources and… See the full description on the dataset page: https://huggingface.co/datasets/MR-CODESPIKE/sentinelng-data-nlu.text0 likes46 downloads19d agoHugging Face11deutsche-telekom /NLU-Evaluation-Data-en-de NLU Evaluation Data - English and German A labeled English and German language multi-domain dataset (21 domains) with 25K user utterances for human-robot interaction. This dataset is collected and annotated for evaluating NLU services and platforms. The detailed paper on this dataset can be found at arXiv.org: Benchmarking Natural Language Understanding Services for building Conversational Agents The dataset builds on the annotated data of the xliuhw/NLU-Evaluation-Data repository.… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-Evaluation-Data-en-de.tabulartext-classification10K<n<100K2 likes44 downloads3y agoHugging Face12aveni-ai /aveni-bench-polyai-nlu AveniBench: PolyAI NLU++ PolyAI NLU++ split used in the AveniBench. License This dataset is made available under the CC-BY-4.0 license. Citation AveniBench TDB PolyAI NLU++ @inproceedings{casanueva-etal-2022-nlu, title = "{NLU}++: A Multi-Label, Slot-Rich, Generalisable Dataset for Natural Language Understanding in Task-Oriented Dialogue", author = "Casanueva, Inigo and Vuli{\'c}, Ivan and Spithourakis, Georgios and Budzianowski… See the full description on the dataset page: https://huggingface.co/datasets/aveni-ai/aveni-bench-polyai-nlu.textn<1K0 likes44 downloads2y agoHugging Face13alfredmh /Palette_nlu_service Persian Sales NLU (seed) Synthetic + templated Persian sales dialogues for NLU training. Splits: train.jsonl, val.jsonl, test.jsonl Each line: text, tokens, intent, slots (BIO), keep_mask, search_query. texttext-classification1K<n<10K0 likes44 downloads4d agoHugging Face14marcuschill1823 /incar-nlu-judge-out-v31text0 likes38 downloads3mo agoHugging Face15KIEFERSA /greek-nlu-bench Greek NLU Benchmark Frozen (v1.0) Greek language-understanding benchmark — 9,751 items / 18 tasks, one split per task. Native-verified and train/eval-disjoint. Built to detect CPT/SFT gains and regressions on Greek, weighted toward morphology-sensitive tasks. Item schema One JSON object per line; identical envelope across tasks: { "id": "el-mcq-000042", "task": "mcq", "language": "el", "version": "1.0", "tags": {"category": "linguistic", "phenomenon":… See the full description on the dataset page: https://huggingface.co/datasets/KIEFERSA/greek-nlu-bench.text1K<n<10K0 likes36 downloads3mo agoHugging Face16deutsche-telekom /NLU-few-shot-benchmark-en-de NLU Few-shot Benchmark - English and German This is a few-shot training dataset from the domain of human-robot interaction. It contains texts in German and English language with 64 different utterances (classes). Each utterance (class) has exactly 20 samples in the training set. This leads to a total of 1280 different training samples. The dataset is intended to benchmark the intent classifiers of chat bots in English and especially in German language. We are building on our… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/NLU-few-shot-benchmark-en-de.texttext-classification10K<n<100K4 likes33 downloads3y agoHugging Face17Palettetech /Palette_nlu_service Persian Sales NLU (seed) Synthetic + templated Persian sales dialogues for NLU training. Splits: train.jsonl, val.jsonl, test.jsonl Each line: text, tokens, intent, slots (BIO), keep_mask, search_query. texttext-classification1K<n<10K0 likes31 downloads3d agoHugging Face18abilfad /indo-nlu-entailmen Dataset Card for "indo-nlu-entailmen" More Information needed textn<1K0 likes25 downloads3y agoHugging Face19nluo421 /Chess_piecesimagen<1K0 likes24 downloads2y agoHugging Face20quocanh34 /new_nlu_tts3 Dataset Card for "new_nlu_tts3" More Information needed text1K<n<10K0 likes22 downloads3y agoHugging Face21nluai /dataset_VMLU_for_bloomtextn<1K0 likes22 downloads2y agoHugging Face22Solmazp /parsi-nlutext1K<n<10K1 likes20 downloads1y agoHugging Face23darkmatter2222 /NLU-Redact-PII-v1 Synthetic Dataset Data Card This document provides an overview of the synthetic dataset generated for testing redaction and anonymization pipelines. It outlines the data generation process, the variety in data formats, ethical considerations, and the impact of complex invalid formats on model quality. Overview The synthetic dataset is created using a suite of generators that produce both valid and intentionally invalid formats for sensitive data such as names, card… See the full description on the dataset page: https://huggingface.co/datasets/darkmatter2222/NLU-Redact-PII-v1.text1K<n<10K0 likes19 downloads2y agoHugging Face24nluai /dataset_dhnl_qna_v2text1K<n<10K1 likes18 downloads2y agoHugging Face25nlux /audiologyaudion<1K0 likes17 downloads2y agoHugging Face26Wikit /nlu-covidFrench benchmark of NLU services for employee support use case during covid-19 pandemic. These datasets were created by the Wikit team in order to compare the performances of NLU tools on the French language. The dataset use case is employee support during the covid 19 pandemic. The intents were defined to answer department employees' questions on the evolution of work conditions related to the crisis. The training_dataset.csv file contains training utterances with associated intent used to… See the full description on the dataset page: https://huggingface.co/datasets/Wikit/nlu-covid.texttext-classification1K<n<10K0 likes16 downloads3y agoHugging Face27mteb /ru_nlu_intenttext10K<n<100K0 likes15 downloads1y agoHugging Face28vsarathy /DIARC-embodied-nlu-styled-4k DIARC-LLM-Parser-Embodied-NLU-Styled-4K This dataset contains about ~4k utterances together with their semantic parses as interpretable by the DIARC cognitive robotic architecture. The parses are meant to capture the speech-theoretic aspects of NL and parse the intent, referents, and descriptors in the utterance. This dataset is one in a set of datasets. For this particular one, we programmatically built 127 utterances and semantics that are groundable in a robotic architecture… See the full description on the dataset page: https://huggingface.co/datasets/vsarathy/DIARC-embodied-nlu-styled-4k.text1K<n<10K2 likes12 downloads3y agoHugging Face29hongdijk /kor_nlu_hufstabular1K<n<10K0 likes11 downloads4y agoHugging Face30nluai /nlu_stsv_testtextn<1K0 likes11 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.