CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /amazon_massive_intent MassiveIntentClassification An MTEB dataset Massive Text Embedding Benchmark MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages Task category t2c Domains Spoken Reference https://arxiv.org/abs/2204.08582 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["MassiveIntentClassification"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/amazon_massive_intent.texttext-classification100K<n<1M27 likes38k downloads7mo agoHugging Face02SetFit /amazon_massive_intent_en-UStext10K<n<100K10 likes1.5k downloads4y agoHugging Face03HugoGiddins /buy_sell_intenttabularn<1K0 likes1.1k downloads2mo agoHugging Face04sonos-nlu-benchmark /snips_built_in_intents Dataset Card for Snips Built In Intents Dataset Summary Snips' built in intents dataset was initially used to compare different voice assistants and released as a public dataset hosted at https://github.com/sonos/nlu-benchmark in folder 2016-12-built-in-intents. The dataset contains 328 utterances over 10 intent classes. A related Medium post is https://medium.com/snips-ai/benchmarking-natural-language-understanding-systems-d35be6ce568d. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sonos-nlu-benchmark/snips_built_in_intents.texttext-classificationn<1K14 likes1k downloads2y agoHugging Face05OpenVoiceOS /ovos-tts-bench-intents-for-eval-prompts OVOS tts bench — intents-for-eval-prompts Synthesised clips (one per prompt) predictions of the registered OVOS Plugin Arena tts fighters over OpenVoiceOS/intents-for-eval. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-tts-bench-intents-for-eval-prompts.audio0 likes768 downloads13d agoHugging Face06hamedrahimi /IntentQAtabular10K<n<100K0 likes757 downloads8mo agoHugging Face07mteb /mtop_intent MTOPIntentClassification An MTEB dataset Massive Text Embedding Benchmark MTOP: Multilingual Task-Oriented Semantic Parsing Task category t2c Domains Spoken, Spoken Reference https://arxiv.org/pdf/2008.09335.pdf How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["MTOPIntentClassification"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/mtop_intent.texttext-classification3 likes703 downloads1y agoHugging Face08CserDu123 /IntentionDataset0 likes690 downloads4mo agoHugging Face09yuweiyin /IntentGrasp IntentGrasp: A Comprehensive Benchmark for Intent Understanding Paper: https://arxiv.org/abs/2605.06832 Authors: Yuwei Yin, Chuyuan Li, Giuseppe Carenini Institute: UBC NLP Group, Department of Computer Science, University of British Columbia Keywords: Intent Understanding, Dataset, Benchmark, LLM, Evaluation, Intentional Fine-Tuning Abstract: Accurately understanding the intent behind speech, conversation, and writing is crucial to the development of helpful Large Language… See the full description on the dataset page: https://huggingface.co/datasets/yuweiyin/IntentGrasp.textquestion-answering100K<n<1M6 likes589 downloads2mo agoHugging Face10ethicalabs /Research-Intent-Judge Research Intent — LLM-as-Judge ▶️ Watch the Video LLM-as-Judge annotations for research paper intent classification, collected through the Echo-DSRN collaborative platform during the OpenAIRE AI Hackathon 2026. The dataset has one split per judge model (Gemma_4_E4B_it_GGUF, Qwen3.6_35B_A3B_GGUF, Bonsai_8B_gguf, ...) plus a human_annotations split with curator annotations. Split names use underscores in place of the dashes in model names (HF does not allow dashes in split… See the full description on the dataset page: https://huggingface.co/datasets/ethicalabs/Research-Intent-Judge.texttext-classification100K<n<1M0 likes471 downloads28d agoHugging Face11SetFit /amazon_massive_intent_zh-CNtext10K<n<100K7 likes428 downloads4y agoHugging Face12SetFit /amazon_massive_intent_de-DEtext10K<n<100K0 likes400 downloads4y agoHugging Face13wendell0218 /Movens-Intent&nbsp;OmniHM-Intent An omni-modal benchmark for evaluating intent-to-humanoid motion generation Evaluation-only split. OmniHM-Intent contains fixed benchmark samples selected from the OmniHM training-data pool. It is intended to evaluate models that were not trained on these exact samples. Any overlap with a model's training data must be disclosed. Overview OmniHM-Intent evaluates whether a humanoid motion model can follow intent expressed through four input… See the full description on the dataset page: https://huggingface.co/datasets/wendell0218/Movens-Intent.audiotext-to-video1K<n<10K0 likes351 downloads20d agoHugging Face14Obaraqreceh /forge-intentdata Forge Intent Dataset Version: 1.0.0 textn<1K2 likes331 downloads19h agoHugging Face15OpenVoiceOS /ovos-localize-intents OpenVoiceOS Localize — Intent Classification Dataset Multilingual intent classification corpus exported from OpenVoiceOS/ovos-localize. Each row is a single expanded utterance labelled with the OVOS skill and intent file that produced it. Templates are fully expanded (bracket alternation resolved); {slot_name} placeholders from .intent files are kept verbatim so models can learn the slot-carrying pattern. Schema Column Description lang BCP-47 locale… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-localize-intents.text-classification100K<n<1M0 likes309 downloads10h agoHugging Face16SetFit /amazon_massive_intent_am-ETtext10K<n<100K1 likes307 downloads4y agoHugging Face17SetFit /amazon_massive_intent_hi-INtext10K<n<100K0 likes287 downloads4y agoHugging Face18peakji /peak-intent-50text100K<n<1M0 likes286 downloads2y agoHugging Face19SetFit /amazon_massive_intent_sw-KEtext10K<n<100K1 likes268 downloads4y agoHugging Face20SetFit /amazon_massive_intent_ar-SAtext10K<n<100K1 likes263 downloads4y agoHugging Face21SetFit /amazon_massive_intent_ja-JPtext10K<n<100K0 likes262 downloads4y agoHugging Face22ScortonAI /online-shoppers-intentiontabular10K<n<100K0 likes243 downloads3y agoHugging Face23OpenVoiceOS /intents-for-eval Purpose. This dataset was collected specifically for intent-parser benchmarking, independently from any OVOS skill. Skill-derived utterances tend to overfit the exact phrasings a plugin was tuned on; this data is drawn from a disjoint source so it measures whether an OVOS intent plugin generalizes rather than memorizes. It is part of the OVOS intent-classification datasets used by the OVOS Plugin Arena intent benchmark. Funding Developed by TigreGotico for OpenVoiceOS as part… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/intents-for-eval.texttext-classification10K<n<100K0 likes225 downloads1mo agoHugging Face24SetFit /amazon_massive_intent_th-THtext10K<n<100K0 likes221 downloads4y agoHugging Face25SetFit /amazon_massive_intent_fr-FRtext10K<n<100K0 likes217 downloads2y agoHugging Face26SetFit /amazon_massive_intent_es-EStext10K<n<100K0 likes211 downloads4y agoHugging Face27SetFit /amazon_massive_intent_my-MMtext10K<n<100K0 likes207 downloads4y agoHugging Face28SetFit /amazon_massive_intent_ru-RUtext10K<n<100K1 likes203 downloads4y agoHugging Face29OpenVoiceOS /MT-intents-dataset-pt-PTtext10K<n<100K0 likes203 downloads1y agoHugging Face30SetFit /amazon_massive_intent_ko-KRtext10K<n<100K0 likes199 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.