CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HugoGiddins /buy_sell_intenttabularn<1K0 likes1.1k downloads2mo agoHugging Face02hamedrahimi /IntentQAtabular10K<n<100K0 likes757 downloads8mo agoHugging Face03wendell0218 /Movens-Intent&nbsp;OmniHM-Intent An omni-modal benchmark for evaluating intent-to-humanoid motion generation Evaluation-only split. OmniHM-Intent contains fixed benchmark samples selected from the OmniHM training-data pool. It is intended to evaluate models that were not trained on these exact samples. Any overlap with a model's training data must be disclosed. Overview OmniHM-Intent evaluates whether a humanoid motion model can follow intent expressed through four input… See the full description on the dataset page: https://huggingface.co/datasets/wendell0218/Movens-Intent.audiotext-to-video1K<n<10K0 likes351 downloads21d agoHugging Face04ScortonAI /online-shoppers-intentiontabular10K<n<100K0 likes243 downloads3y agoHugging Face05OpenVoiceOS /ovos-intent-bench-intents-for-eval OVOS intent bench — intents-for-eval Per-sample predictions of the open intent league (mixed-paradigm pipeline fusions) fighters of the OVOS Plugin Arena over OpenVoiceOS/intents-for-eval. One dedicated repo per benchmark modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, fired pipeline stage, exact_match with correct-OOD semantics).… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-intents-for-eval.tabular100K<n<1M0 likes196 downloads16d agoHugging Face06OpenVoiceOS /ovos-intent-keyword-bench-intents-for-eval OVOS intent_keyword bench — intents-for-eval Per-sample predictions of the keyword-paradigm intent league fighters of the OVOS Plugin Arena over OpenVoiceOS/intents-for-eval. One dedicated repo per benchmark modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, fired pipeline stage, exact_match with correct-OOD semantics). Produced by the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-keyword-bench-intents-for-eval.tabular10K<n<100K0 likes170 downloads16d agoHugging Face07Jazhyc /aims-safety-intents AIMS: Annotated Intents for Model Safety AIMS is a human-annotated dataset of user intents for LLM safety classification. Each example pairs a difficult safety prompt with a concise, human-written description of the user's underlying intent and a human-assigned harm label. The dataset is built to study a single question: can safety classifiers be improved by modeling why a user is asking something, rather than relying on surface-level text cues? It contains 1,724 annotated… See the full description on the dataset page: https://huggingface.co/datasets/Jazhyc/aims-safety-intents.tabulartext-classification1K<n<10K2 likes85 downloads25d agoHugging Face08nraptisss /TMF921-intent-to-config-25k TMF921-Grounded Intent-to-Network-Configuration Dataset (25K) The most comprehensive open dataset for training LLMs to translate natural language network intents into spec-compliant 5G/6G configurations. 25,000 samples (22,500 train / 2,500 test) of natural language intents paired with structured network configurations across 6 target specification layers and 8 lifecycle operations, all grounded in real telecom standards. What Makes This Dataset Unique… See the full description on the dataset page: https://huggingface.co/datasets/nraptisss/TMF921-intent-to-config-25k.documenttext-generation10K<n<100K0 likes75 downloads5mo agoHugging Face09OpenVoiceOS /ovos-intent-bench-speech-massive-nl-NL ovos-intent-bench-speech-massive-nl-NL An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-nl-NL dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-nl-NL.tabularn<1K0 likes70 downloads18d agoHugging Face10OpenVoiceOS /ovos-intent-bench-speech-massive-pt-PT ovos-intent-bench-speech-massive-pt-PT An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-pt-PT dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-pt-PT.tabularn<1K0 likes70 downloads18d agoHugging Face11OpenVoiceOS /ovos-intent-bench-speech-massive-ru-RU ovos-intent-bench-speech-massive-ru-RU An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-ru-RU dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-ru-RU.tabularn<1K0 likes70 downloads18d agoHugging Face12PrevenIA /spanish-suicide-intent Dataset Summary The dataset consists of comments from several sources translated to Spanish language and classified as suicidal ideation/behavior and non-suicidal. Dataset Structure The dataset has 175010 rows (77223 considered as Suicidal Ideation/Behavior and 97787 considered Not Suicidal). Dataset fields Text: User comment. Label: 1 if suicidal ideation/behavior; 0 if not suicidal comment. Dataset: Source of the comment Dataset Creation 112385… See the full description on the dataset page: https://huggingface.co/datasets/PrevenIA/spanish-suicide-intent.tabulartext-classification100K<n<1M3 likes69 downloads3y agoHugging Face13OpenVoiceOS /ovos-intent-bench-speech-massive-pl-PL ovos-intent-bench-speech-massive-pl-PL An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-pl-PL dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-pl-PL.tabularn<1K0 likes69 downloads18d agoHugging Face14OpenVoiceOS /ovos-intent-bench-speech-massive-ar-SA ovos-intent-bench-speech-massive-ar-SA An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-ar-SA dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-ar-SA.tabularn<1K0 likes68 downloads18d agoHugging Face15OpenVoiceOS /ovos-intent-bench-speech-massive-de-DE ovos-intent-bench-speech-massive-de-DE An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-de-DE dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-de-DE.tabularn<1K0 likes68 downloads18d agoHugging Face16OpenVoiceOS /ovos-intent-bench-speech-massive-fr-FR ovos-intent-bench-speech-massive-fr-FR An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-fr-FR dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-fr-FR.tabularn<1K0 likes68 downloads18d agoHugging Face17OpenVoiceOS /ovos-intent-bench-speech-massive-hu-HU ovos-intent-bench-speech-massive-hu-HU An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-hu-HU dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-hu-HU.tabularn<1K0 likes67 downloads18d agoHugging Face18OpenVoiceOS /ovos-intent-bench-speech-massive-vi-VN ovos-intent-bench-speech-massive-vi-VN An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-vi-VN dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-vi-VN.tabularn<1K0 likes67 downloads18d agoHugging Face19nraptisss /TMF921-intent-to-config-research-sota TMF921 Intent-to-Config Research SOTA Splits This dataset is a research-oriented derivative of nraptisss/TMF921-intent-to-config-augmented. It provides reproducible training and OOD evaluation splits for supervised fine-tuning models that translate natural-language telecom/network-slicing intents into structured JSON configuration objects. This dataset is intended for research. It is not a production-certified telecom configuration generator and should not be used to deploy network… See the full description on the dataset page: https://huggingface.co/datasets/nraptisss/TMF921-intent-to-config-research-sota.tabulartext-generation10K<n<100K0 likes66 downloads5mo agoHugging Face20OpenVoiceOS /ovos-intent-bench-speech-massive-es-ES ovos-intent-bench-speech-massive-es-ES An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-es-ES dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-es-ES.tabularn<1K0 likes63 downloads18d agoHugging Face21OpenVoiceOS /ovos-intent-bench-speech-massive-ko-KR ovos-intent-bench-speech-massive-ko-KR An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-ko-KR dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-ko-KR.tabularn<1K0 likes63 downloads18d agoHugging Face22nraptisss /telecom-intent-config-sft-10k Telecom Intent→Config SFT Dataset (10K) The first open SFT dataset for training LLMs to translate natural language network intents into structured 5G/6G configurations. This dataset addresses the #1 gap identified in the telecom LLM research landscape: there is no public training dataset for intent-to-policy translation. All existing telecom datasets (TeleQnA, ORANBench-13K, 6G-Bench) are MCQ evaluation benchmarks — not instruction-following format. This dataset fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/nraptisss/telecom-intent-config-sft-10k.tabulartext-generation10K<n<100K1 likes51 downloads5mo agoHugging Face23avsolatorio /mteb-mtop_intent-avs_triplets MTEB MTOP Intent Triplets Dataset This dataset was used in the paper GISTEmbed: Guided In-sample Selection of Training Negatives for Text Embedding Fine-tuning. Refer to https://arxiv.org/abs/2402.16829 for details. The code for generating the data is available at https://github.com/avsolatorio/GISTEmbed/blob/main/scripts/create_classification_dataset.py. Citation @article{solatorio2024gistembed, title={GISTEmbed: Guided In-sample Selection of Training Negatives… See the full description on the dataset page: https://huggingface.co/datasets/avsolatorio/mteb-mtop_intent-avs_triplets.tabular10K<n<100K0 likes49 downloads3y agoHugging Face24sempite /product-tagging-saved-intent Product Tagging and Saved Intent in Online Retail How online stores structure product tags and implement saved-item features, and how that compares with a platform where tagging is crowd-sourced against a shared controlled vocabulary. Canonical release: https://doi.org/10.5281/zenodo.22852049 This repository mirrors that deposit. Cite the DOI. Sample 1,028 Shopify storefronts, drawn by seeded random sample from a 30,000-domain draw of the Tranco top 1M, measured… See the full description on the dataset page: https://huggingface.co/datasets/sempite/product-tagging-saved-intent.tabular10K<n<100K0 likes46 downloads2d agoHugging Face25OnMoon2 /intent-classifier-v4tabular100K<n<1M0 likes45 downloads6mo agoHugging Face26rungalileo /banking_intenttabular10K<n<100K2 likes40 downloads4y agoHugging Face27avsolatorio /mteb-amazon_massive_intent-avs_triplets MTEB Amazon Massive Intent Triplets Dataset This dataset was used in the paper GISTEmbed: Guided In-sample Selection of Training Negatives for Text Embedding Fine-tuning. Refer to https://arxiv.org/abs/2402.16829 for details. The code for generating the data is available at https://github.com/avsolatorio/GISTEmbed/blob/main/scripts/create_classification_dataset.py. Citation @article{solatorio2024gistembed, title={GISTEmbed: Guided In-sample Selection of Training… See the full description on the dataset page: https://huggingface.co/datasets/avsolatorio/mteb-amazon_massive_intent-avs_triplets.tabular10K<n<100K0 likes39 downloads3y agoHugging Face28CharlieJi /HelpSteer2_with_intent Dataset Card for HelpSteer2_with_intent This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/CharlieJi/HelpSteer2_with_intent/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/CharlieJi/HelpSteer2_with_intent.tabularn<1K0 likes38 downloads2y agoHugging Face29ClarusC64 /autonomous-driving-intention-field-extraction-v0.1What this dataset tests Whether a system can infer agent intentions from context cues in complex driving scenes. This is not trajectory prediction. It is intention inference. Required outputs agent_id inferred_intention intention_confidence time_horizon_s alternative_intentions stability_score Scoring conventions confidence and stability range 0 to 1 time horizon is seconds into the near future Use case Layer one of Intention Field and Social Coherence Maps. This enables… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-intention-field-extraction-v0.1.tabulartabular-classificationn<1K0 likes37 downloads8mo agoHugging Face30DFKI /radr_intents Dataset Card for "Intent Classification for Robot Assisted Disaster Response" This dataset consists of conversations recorded during the training sessions in the emergency response domain. The conversations are typically between several operators controlling the robots, a team leader and a mission commander. The data have been transcribed and annotated during the following projects: TRADR and ADRZ. The dialogues are split into turns and each turn is annotated with a speaker and… See the full description on the dataset page: https://huggingface.co/datasets/DFKI/radr_intents.tabulartext-classification1K<n<10K0 likes34 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.