datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amazon_massive_intent
MassiveIntentClassification
An MTEB dataset
Massive Text Embedding Benchmark
MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages
Task category
t2c
Domains
Spoken
Reference
https://arxiv.org/abs/2204.08582
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["MassiveIntentClassification"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/amazon_massive_intent.amazon_massive_intent_en-USbuy_sell_intentsnips_built_in_intents
Dataset Card for Snips Built In Intents
Dataset Summary
Snips' built in intents dataset was initially used to compare different voice assistants and released as a public dataset hosted at
https://github.com/sonos/nlu-benchmark in folder 2016-12-built-in-intents. The dataset contains 328 utterances over 10 intent classes.
A related Medium post is https://medium.com/snips-ai/benchmarking-natural-language-understanding-systems-d35be6ce568d.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sonos-nlu-benchmark/snips_built_in_intents.ovos-tts-bench-intents-for-eval-prompts
OVOS tts bench — intents-for-eval-prompts
Synthesised clips (one per prompt) predictions of the registered
OVOS Plugin Arena
tts fighters over
OpenVoiceOS/intents-for-eval.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-tts-bench-intents-for-eval-prompts.IntentQAmtop_intent
MTOPIntentClassification
An MTEB dataset
Massive Text Embedding Benchmark
MTOP: Multilingual Task-Oriented Semantic Parsing
Task category
t2c
Domains
Spoken, Spoken
Reference
https://arxiv.org/pdf/2008.09335.pdf
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["MTOPIntentClassification"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/mtop_intent.IntentionDatasetIntentGrasp
IntentGrasp: A Comprehensive Benchmark for Intent Understanding
Paper: https://arxiv.org/abs/2605.06832
Authors: Yuwei Yin, Chuyuan Li, Giuseppe Carenini
Institute: UBC NLP Group, Department of Computer Science, University of British Columbia
Keywords: Intent Understanding, Dataset, Benchmark, LLM, Evaluation, Intentional Fine-Tuning
Abstract:
Accurately understanding the intent behind speech, conversation, and writing is crucial to the development of helpful Large Language… See the full description on the dataset page: https://huggingface.co/datasets/yuweiyin/IntentGrasp.Research-Intent-Judge
Research Intent — LLM-as-Judge
▶️ Watch the Video
LLM-as-Judge annotations for research paper intent classification, collected
through the Echo-DSRN collaborative platform during the OpenAIRE AI Hackathon 2026.
The dataset has one split per judge model (Gemma_4_E4B_it_GGUF,
Qwen3.6_35B_A3B_GGUF, Bonsai_8B_gguf, ...) plus a human_annotations
split with curator annotations. Split names use underscores in place of the
dashes in model names (HF does not allow dashes in split… See the full description on the dataset page: https://huggingface.co/datasets/ethicalabs/Research-Intent-Judge.amazon_massive_intent_zh-CNamazon_massive_intent_de-DEMovens-Intent OmniHM-Intent
An omni-modal benchmark for evaluating intent-to-humanoid motion generation
Evaluation-only split. OmniHM-Intent contains fixed benchmark samples selected from the OmniHM training-data pool. It is intended to evaluate models that were not trained on these exact samples. Any overlap with a model's training data must be disclosed.
Overview
OmniHM-Intent evaluates whether a humanoid motion model can follow intent expressed through four input… See the full description on the dataset page: https://huggingface.co/datasets/wendell0218/Movens-Intent.forge-intentdata
Forge Intent Dataset
Version: 1.0.0
ovos-localize-intents
OpenVoiceOS Localize — Intent Classification Dataset
Multilingual intent classification corpus exported from
OpenVoiceOS/ovos-localize.
Each row is a single expanded utterance labelled with the OVOS skill and
intent file that produced it. Templates are fully expanded (bracket
alternation resolved); {slot_name} placeholders from .intent files are
kept verbatim so models can learn the slot-carrying pattern.
Schema
Column
Description
lang
BCP-47 locale… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-localize-intents.amazon_massive_intent_am-ETamazon_massive_intent_hi-INpeak-intent-50amazon_massive_intent_sw-KEamazon_massive_intent_ar-SAamazon_massive_intent_ja-JPonline-shoppers-intentionintents-for-eval
Purpose. This dataset was collected specifically for intent-parser benchmarking, independently from any OVOS skill. Skill-derived utterances tend to overfit the exact phrasings a plugin was tuned on; this data is drawn from a disjoint source so it measures whether an OVOS intent plugin generalizes rather than memorizes. It is part of the OVOS intent-classification datasets used by the OVOS Plugin Arena intent benchmark.
Funding
Developed by TigreGotico for OpenVoiceOS as part… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/intents-for-eval.amazon_massive_intent_th-THamazon_massive_intent_fr-FRamazon_massive_intent_es-ESamazon_massive_intent_my-MMamazon_massive_intent_ru-RUMT-intents-dataset-pt-PTamazon_massive_intent_ko-KR
