datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ovos-tts-bench-intents-for-eval-prompts
OVOS tts bench — intents-for-eval-prompts
Synthesised clips (one per prompt) predictions of the registered
OVOS Plugin Arena
tts fighters over
OpenVoiceOS/intents-for-eval.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-tts-bench-intents-for-eval-prompts.Movens-Intent OmniHM-Intent
An omni-modal benchmark for evaluating intent-to-humanoid motion generation
Evaluation-only split. OmniHM-Intent contains fixed benchmark samples selected from the OmniHM training-data pool. It is intended to evaluate models that were not trained on these exact samples. Any overlap with a model's training data must be disclosed.
Overview
OmniHM-Intent evaluates whether a humanoid motion model can follow intent expressed through four input… See the full description on the dataset page: https://huggingface.co/datasets/wendell0218/Movens-Intent.slurp_slu_intentcall-transcript-intent-data-v2
Call Transcript Intent Dataset
Multimodal Hindi/Hinglish customer utterance dataset for loan/EMI/payment call intent classification.
Dataset Summary
Metric
Value
Total examples
139,348
Total audio duration
51.04 h
Number of intents
17
Split Statistics
Split
Examples
Duration
Hours
train
126,848
2755.14 min
45.92 h
validation
10,000
219.03 min
3.65 h
eval
2,500
88.37 min
1.47 h
Class Distribution… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/call-transcript-intent-data-v2.medical-intent-audio-datasetslurp_slu_intent_with_transcriptionMedical_Speech_Transcription_and_IntentThis dataset came from Kaggle and was contributed by Paul Mooney.
https://www.kaggle.com/datasets/paultimothymooney/medical-speech-transcription-and-intent/data
Context
8.5 hours of audio utterances paired with text for common medical symptoms.
Content
This data contains thousands of audio utterances for common medical symptoms like “knee pain” or “headache,” totaling more than 8 hours in aggregate. Each utterance was created by individual human contributors based on a given symptom. These… See the full description on the dataset page: https://huggingface.co/datasets/Shamus/Medical_Speech_Transcription_and_Intent.IntentTrain
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
This repository contains the dataset and associated information for the paper HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.
👀 HumanOmniV2 Overview
With the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability, which demands detailed and thoughtful reasoning. In recent studies… See the full description on the dataset page: https://huggingface.co/datasets/PhilipC/IntentTrain.medical-intent-audio-dataset-consolidatedSuperbIC_SLURP-Intentspeech-to-intent-benchmark
Speech-to-Intent Benchmark
Made in Vancouver, Canada by Picovoice
This framework benchmarks the accuracy of Picovoice's Speech-to-Intent engine, Rhino.
It compares the accuracy of Rhino with:
Amazon Lex
Google Dialogflow
IBM Watson
Microsoft LUIS
Results
Command acceptance rate is the probability of an engine correctly understanding the spoken command. Below is the summary:
The figure below depicts engines performance at each SNR:
Data
The speech data… See the full description on the dataset page: https://huggingface.co/datasets/Picovoice/speech-to-intent-benchmark.intent-datasetIntentClassification_FluentSpeechCommands-Action_TTSIntent_ClassificationIntentClassification_FluentSpeechCommands-Action_TTSIntentClassification_FluentSpeechCommands-Location_TTSIntentClassification_FluentSpeechCommands-Object_TTSIntentClassification_FluentSpeechCommands-Action
