CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /amazon_massive_intent MassiveIntentClassification An MTEB dataset Massive Text Embedding Benchmark MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages Task category t2c Domains Spoken Reference https://arxiv.org/abs/2204.08582 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["MassiveIntentClassification"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/amazon_massive_intent.texttext-classification100K<n<1M27 likes38k downloads7mo agoHugging Face02SetFit /amazon_massive_intent_en-UStext10K<n<100K10 likes1.5k downloads4y agoHugging Face03HugoGiddins /buy_sell_intenttabularn<1K0 likes1.1k downloads2mo agoHugging Face04sonos-nlu-benchmark /snips_built_in_intents Dataset Card for Snips Built In Intents Dataset Summary Snips' built in intents dataset was initially used to compare different voice assistants and released as a public dataset hosted at https://github.com/sonos/nlu-benchmark in folder 2016-12-built-in-intents. The dataset contains 328 utterances over 10 intent classes. A related Medium post is https://medium.com/snips-ai/benchmarking-natural-language-understanding-systems-d35be6ce568d. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sonos-nlu-benchmark/snips_built_in_intents.texttext-classificationn<1K14 likes1.1k downloads2y agoHugging Face05hamedrahimi /IntentQAtabular10K<n<100K0 likes757 downloads8mo agoHugging Face06mteb /mtop_intent MTOPIntentClassification An MTEB dataset Massive Text Embedding Benchmark MTOP: Multilingual Task-Oriented Semantic Parsing Task category t2c Domains Spoken, Spoken Reference https://arxiv.org/pdf/2008.09335.pdf How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["MTOPIntentClassification"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/mtop_intent.texttext-classification3 likes711 downloads1y agoHugging Face07yuweiyin /IntentGrasp IntentGrasp: A Comprehensive Benchmark for Intent Understanding Paper: https://arxiv.org/abs/2605.06832 Authors: Yuwei Yin, Chuyuan Li, Giuseppe Carenini Institute: UBC NLP Group, Department of Computer Science, University of British Columbia Keywords: Intent Understanding, Dataset, Benchmark, LLM, Evaluation, Intentional Fine-Tuning Abstract: Accurately understanding the intent behind speech, conversation, and writing is crucial to the development of helpful Large Language… See the full description on the dataset page: https://huggingface.co/datasets/yuweiyin/IntentGrasp.textquestion-answering100K<n<1M6 likes596 downloads2mo agoHugging Face08ethicalabs /Research-Intent-Judge Research Intent — LLM-as-Judge ▶️ Watch the Video LLM-as-Judge annotations for research paper intent classification, collected through the Echo-DSRN collaborative platform during the OpenAIRE AI Hackathon 2026. The dataset has one split per judge model (Gemma_4_E4B_it_GGUF, Qwen3.6_35B_A3B_GGUF, Bonsai_8B_gguf, ...) plus a human_annotations split with curator annotations. Split names use underscores in place of the dashes in model names (HF does not allow dashes in split… See the full description on the dataset page: https://huggingface.co/datasets/ethicalabs/Research-Intent-Judge.texttext-classification100K<n<1M0 likes470 downloads29d agoHugging Face09SetFit /amazon_massive_intent_zh-CNtext10K<n<100K7 likes448 downloads4y agoHugging Face10SetFit /amazon_massive_intent_de-DEtext10K<n<100K0 likes401 downloads4y agoHugging Face11wendell0218 /Movens-Intent&nbsp;OmniHM-Intent An omni-modal benchmark for evaluating intent-to-humanoid motion generation Evaluation-only split. OmniHM-Intent contains fixed benchmark samples selected from the OmniHM training-data pool. It is intended to evaluate models that were not trained on these exact samples. Any overlap with a model's training data must be disclosed. Overview OmniHM-Intent evaluates whether a humanoid motion model can follow intent expressed through four input… See the full description on the dataset page: https://huggingface.co/datasets/wendell0218/Movens-Intent.audiotext-to-video1K<n<10K0 likes370 downloads21d agoHugging Face12Obaraqreceh /forge-intentdata Forge Intent Dataset Version: 1.0.0 textn<1K2 likes331 downloads5h agoHugging Face13SetFit /amazon_massive_intent_am-ETtext10K<n<100K1 likes308 downloads4y agoHugging Face14peakji /peak-intent-50text100K<n<1M0 likes296 downloads2y agoHugging Face15SetFit /amazon_massive_intent_hi-INtext10K<n<100K0 likes290 downloads4y agoHugging Face16SetFit /amazon_massive_intent_ja-JPtext10K<n<100K0 likes271 downloads4y agoHugging Face17SetFit /amazon_massive_intent_ar-SAtext10K<n<100K1 likes268 downloads4y agoHugging Face18SetFit /amazon_massive_intent_sw-KEtext10K<n<100K1 likes266 downloads4y agoHugging Face19ScortonAI /online-shoppers-intentiontabular10K<n<100K0 likes252 downloads3y agoHugging Face20SetFit /amazon_massive_intent_th-THtext10K<n<100K0 likes224 downloads4y agoHugging Face21OpenVoiceOS /MT-intents-dataset-pt-PTtext10K<n<100K0 likes222 downloads1y agoHugging Face22SetFit /amazon_massive_intent_fr-FRtext10K<n<100K0 likes221 downloads2y agoHugging Face23OpenVoiceOS /intents-for-eval Purpose. This dataset was collected specifically for intent-parser benchmarking, independently from any OVOS skill. Skill-derived utterances tend to overfit the exact phrasings a plugin was tuned on; this data is drawn from a disjoint source so it measures whether an OVOS intent plugin generalizes rather than memorizes. It is part of the OVOS intent-classification datasets used by the OVOS Plugin Arena intent benchmark. Funding Developed by TigreGotico for OpenVoiceOS as part… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/intents-for-eval.texttext-classification10K<n<100K0 likes221 downloads1mo agoHugging Face24SetFit /amazon_massive_intent_es-EStext10K<n<100K0 likes216 downloads4y agoHugging Face25SetFit /amazon_massive_intent_my-MMtext10K<n<100K0 likes207 downloads4y agoHugging Face26SetFit /amazon_massive_intent_ru-RUtext10K<n<100K1 likes203 downloads4y agoHugging Face27SetFit /amazon_massive_intent_ko-KRtext10K<n<100K0 likes202 downloads4y agoHugging Face28OpenVoiceOS /ovos-intent-bench-intents-for-eval OVOS intent bench — intents-for-eval Per-sample predictions of the open intent league (mixed-paradigm pipeline fusions) fighters of the OVOS Plugin Arena over OpenVoiceOS/intents-for-eval. One dedicated repo per benchmark modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, fired pipeline stage, exact_match with correct-OOD semantics).… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-intents-for-eval.tabular100K<n<1M0 likes198 downloads16d agoHugging Face29SetFit /amazon_massive_intent_fa-IRtext10K<n<100K0 likes197 downloads4y agoHugging Face30SetFit /amazon_massive_intent_he-ILtext10K<n<100K0 likes193 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.