datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amazon_massive_intent
MassiveIntentClassification
An MTEB dataset
Massive Text Embedding Benchmark
MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages
Task category
t2c
Domains
Spoken
Reference
https://arxiv.org/abs/2204.08582
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["MassiveIntentClassification"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/amazon_massive_intent.amazon_massive_intent_en-USbuy_sell_intentsnips_built_in_intents
Dataset Card for Snips Built In Intents
Dataset Summary
Snips' built in intents dataset was initially used to compare different voice assistants and released as a public dataset hosted at
https://github.com/sonos/nlu-benchmark in folder 2016-12-built-in-intents. The dataset contains 328 utterances over 10 intent classes.
A related Medium post is https://medium.com/snips-ai/benchmarking-natural-language-understanding-systems-d35be6ce568d.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sonos-nlu-benchmark/snips_built_in_intents.IntentQAmtop_intent
MTOPIntentClassification
An MTEB dataset
Massive Text Embedding Benchmark
MTOP: Multilingual Task-Oriented Semantic Parsing
Task category
t2c
Domains
Spoken, Spoken
Reference
https://arxiv.org/pdf/2008.09335.pdf
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["MTOPIntentClassification"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/mtop_intent.IntentGrasp
IntentGrasp: A Comprehensive Benchmark for Intent Understanding
Paper: https://arxiv.org/abs/2605.06832
Authors: Yuwei Yin, Chuyuan Li, Giuseppe Carenini
Institute: UBC NLP Group, Department of Computer Science, University of British Columbia
Keywords: Intent Understanding, Dataset, Benchmark, LLM, Evaluation, Intentional Fine-Tuning
Abstract:
Accurately understanding the intent behind speech, conversation, and writing is crucial to the development of helpful Large Language… See the full description on the dataset page: https://huggingface.co/datasets/yuweiyin/IntentGrasp.Research-Intent-Judge
Research Intent — LLM-as-Judge
▶️ Watch the Video
LLM-as-Judge annotations for research paper intent classification, collected
through the Echo-DSRN collaborative platform during the OpenAIRE AI Hackathon 2026.
The dataset has one split per judge model (Gemma_4_E4B_it_GGUF,
Qwen3.6_35B_A3B_GGUF, Bonsai_8B_gguf, ...) plus a human_annotations
split with curator annotations. Split names use underscores in place of the
dashes in model names (HF does not allow dashes in split… See the full description on the dataset page: https://huggingface.co/datasets/ethicalabs/Research-Intent-Judge.amazon_massive_intent_zh-CNamazon_massive_intent_de-DEMovens-Intent OmniHM-Intent
An omni-modal benchmark for evaluating intent-to-humanoid motion generation
Evaluation-only split. OmniHM-Intent contains fixed benchmark samples selected from the OmniHM training-data pool. It is intended to evaluate models that were not trained on these exact samples. Any overlap with a model's training data must be disclosed.
Overview
OmniHM-Intent evaluates whether a humanoid motion model can follow intent expressed through four input… See the full description on the dataset page: https://huggingface.co/datasets/wendell0218/Movens-Intent.forge-intentdata
Forge Intent Dataset
Version: 1.0.0
amazon_massive_intent_am-ETpeak-intent-50amazon_massive_intent_hi-INamazon_massive_intent_ja-JPamazon_massive_intent_ar-SAamazon_massive_intent_sw-KEonline-shoppers-intentionamazon_massive_intent_th-THMT-intents-dataset-pt-PTamazon_massive_intent_fr-FRintents-for-eval
Purpose. This dataset was collected specifically for intent-parser benchmarking, independently from any OVOS skill. Skill-derived utterances tend to overfit the exact phrasings a plugin was tuned on; this data is drawn from a disjoint source so it measures whether an OVOS intent plugin generalizes rather than memorizes. It is part of the OVOS intent-classification datasets used by the OVOS Plugin Arena intent benchmark.
Funding
Developed by TigreGotico for OpenVoiceOS as part… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/intents-for-eval.amazon_massive_intent_es-ESamazon_massive_intent_my-MMamazon_massive_intent_ru-RUamazon_massive_intent_ko-KRovos-intent-bench-intents-for-eval
OVOS intent bench — intents-for-eval
Per-sample predictions of the open intent league (mixed-paradigm pipeline fusions) fighters of the
OVOS Plugin Arena over
OpenVoiceOS/intents-for-eval.
One dedicated repo per benchmark modality; one dataset split per language;
one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl.
Rows follow the arena §3.2 contract (pinned dataset_revision,
plugin_version, fired pipeline stage, exact_match with correct-OOD
semantics).… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-intents-for-eval.amazon_massive_intent_fa-IRamazon_massive_intent_he-IL
