datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.denoising-impact-evaluation-dataset
Denoising Impact Evaluation Dataset
Dataset Description
The ekacare/denoising-impact-evaluation-dataset is a comprehensive benchmark dataset designed to evaluate the effects of speech enhancement on automatic speech recognition (ASR) systems in medical speech contexts. It includes paired noisy and denoised audio subsets under controlled acoustic conditions to support systematic analysis of denoising performance.
Source Data
Base Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/denoising-impact-evaluation-dataset.Evaluation-Multilingual-VC
Evaluation-Multilingual-VC
We use dataset https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon,
Filter languages that support by Whisper Large V3 to evaluate WER automatically,
Only take test set, sort by up votes.
Because VC required to source text, source audio, target text, we make sure the target text is not same as source text, target text we take from other rows.
Only build first 500 rows for each language
Github issue at… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Evaluation-Multilingual-VC.AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.Stems-Evaluation-Kit
🎧 SonicSets High-Fidelity Stems Evaluation Kit
This is a premium evaluation subset provided by SonicSets, the industrial-grade audio data infrastructure for Large Audio Models (LAM).
📊 Dataset Specifications
Format: 48kHz / 24-bit Uncompressed WAV (Studio-Grade Ground Truth)
Feature: Absolute zero-crosstalk multi-track isolation
Environment: Strict anechoic capture (RT60 < 0.2s)
Purpose: Fully optimized for training and benchmarking state-of-the-art Source… See the full description on the dataset page: https://huggingface.co/datasets/drizzymedia/Stems-Evaluation-Kit.nvidia-brain-noise-evaluation-dataset
Nvidia Brain Noise Evaluation Dataset
Dataset Description
This dataset contains 64 samples organized across multiple splits and 32 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
noisy-bg-snr-10: 2 samples
test: 2 samples
noisy-bg-snr-20: 2 samples
test: 2 samples
noisy-bg-snr-30: 2 samples
test: 2 samples
noisy-bg-snr-40: 2 samples
test: 2 samples
noisy-bg-snr-50: 2 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/nvidia-brain-noise-evaluation-dataset.evaluation-dataset
DeepSafe Evaluation Dataset
Evaluation set for DeepSafe,
a deepfake detection benchmark.
Tiers
Tier
Samples
Generators
Size
Use
master_eval_small/
198
116
1.7 GB
smoke test, under 2 min
master_eval/
15,454
411
10 GB
the standard benchmark
master_eval_full/
45,954
411
25 GB
complete set
Medium tier composition: 9,954 image, 3,500 audio, 2,000 video.
from huggingface_hub import snapshot_download
snapshot_download("deepsafe/evaluation-dataset"… See the full description on the dataset page: https://huggingface.co/datasets/deepsafe/evaluation-dataset.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.telugu-indicf5-evaluationStems-Evaluation-Kit
🎧 SonicSets High-Fidelity Stems Evaluation Kit
This is a premium evaluation subset provided by SonicSets, the industrial-grade audio data infrastructure for Large Audio Models (LAM).
📊 Dataset Specifications
Format: 48kHz / 24-bit Uncompressed WAV (Studio-Grade Ground Truth)
Feature: Absolute zero-crosstalk multi-track isolation
Environment: Strict anechoic capture (RT60 < 0.2s)
Purpose: Fully optimized for training and benchmarking state-of-the-art Source Separation… See the full description on the dataset page: https://huggingface.co/datasets/sonicsets-data/Stems-Evaluation-Kit.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.recorded_evaluation_dataset2denoised-evaluation-dataset
Denoised Subset Fixed
Dataset Description
This dataset contains 500 samples organized across multiple splits and 1 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
denoised: 500 samples
test: 500 samples
Usage
Load specific subset and split:
from datasets import load_dataset
# Load specific subset and split
dataset = load_dataset('sujalappa/denoised-evaluation-dataset'… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/denoised-evaluation-dataset.asr-evaluationsruEnvAQA
ruEnvAQA
Описание задачи
ruEnvAQA – датасет вопросов с множественным и бинарным выбором ответа на русском языке. Вопросы связаны с анализом музыки и невербальных аудиосигналов. Датасет составлен на основе вопросов из англоязычных датасетов Clotho-AQA и MUSIC-AVQA. Вопросы переведены на русский язык и частично изменены, тогда как аудиозаписи использованы в исходном виде (с обрезкой по длине).
Датасет включает вопросы 8 типов:
Оригинальные классы вопросов из MUSIC-AVQA… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/ruEnvAQA.speech-brain-noise-evaluation-dataset
Speech Brain Noise Evaluation Dataset
Dataset Description
This dataset contains 2,000 samples organized across multiple splits and 20 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
noisy-bg-snr-10: 100 samples
test: 100 samples
noisy-bg-snr-30: 100 samples
test: 100 samples
noisy-bg-snr-50: 100 samples
test: 100 samples
denoised-bg-snr-10: 100 samples
test: 100 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/speech-brain-noise-evaluation-dataset.speaker_evaluation_multi_test_v0
Seamless Interaction Pairs
This dataset contains paired query and document audio clips for interaction-based
speaker evaluation. Each row describes a query clip and a related document clip,
with segment metadata and durations for analysis.
Data structure
The dataset uses a single split stored in data.parquet.
Audio files are stored under audio/ and referenced by relative paths in the
parquet file.
Columns
pair_id (string): Pair identifier.
interaction… See the full description on the dataset page: https://huggingface.co/datasets/humanify/speaker_evaluation_multi_test_v0.polyvox-kpi-evaluation
PolyVox KPI Evaluation Dataset
This dataset supports PolyVox evaluation for Problem Statement 11: Real-Time Multi-User Smart Assistant for Dynamic and Noisy Smart Environments.
It contains compact evaluation mixtures for 2-speaker and 3-speaker speech separation under clean, noisy, overlap, and optional RIR-style conditions.
Contents
manifests/kpi_dataset_manifest.csv
dataset_summary.json
audio/ containing mixtures and clean source references
Source… See the full description on the dataset page: https://huggingface.co/datasets/isaacnetero/polyvox-kpi-evaluation.AQUARIA
AQUARIA
Описание задачи
Датасет состоит из вопросов с выбором ответа, проверяющие комплексное понимание аудио, в том числе речи, неречевых сигналов и музыки. Вопросы датасета составлялись таким образом, чтобы для ответа на них требовалось не только распознавать речь, но и анализировать аудиоситуацию целиком и взаимодействие её компонентов. Используемые аудиофайлы созданы специально для датасета AQUARIA.
В датасете представлены вопросы 9 типов:
Audio scene classification… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/AQUARIA.speaker_evaluation_single_test_v0
Speaker Task Test Dataset
数据集描述
40 pairs sampled from Voxceleb1
118 pairs sampled from Voxceleb2
ruSLUn
RuSLUn
Описание задачи
RuSLUn (Russian Spoken Language UNderstanding dataset) — это датасет для задачи понимания устной речи на русском языке, построенный по принципу англоязычного датасета SLURP и мультиязычного xSID, но с учетом культурных и языковых особенностей России. Он предназначен для оценки моделей, которые напрямую преобразуют аудиозаписи в семантическое представление, включая определение намерений пользователя (intent detection) и извлечение слотов (slot… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/ruSLUn.fleurs-evaluationruTiE-Audio
ruTiE-Audio
Описание задачи
ruTiE-Audio — мультимодальная эмуляция теста Тьюринга. Задача сформирована как неизменяемая последовательность вопросно-ответных заданий с опцией выбора ответа. Это 3 связных диалога, каждый с имитацией 500 обращений пользователя к модели. На вход модели подаётся аудио с заключёнными в аудиофайле заданиями и вопросами. Варианты ответа (4 к каждому заданию) модель получает текстом и выбирает из них.
Задания теста проверяют способность модели… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/ruTiE-Audio.tts-evaluation-datasetfleurs-reducedbaseline-model-evaluationsasr_evaluation_datasetsevaluation-whisper-large-v3-floresevaluation-setnepali_asr_evaluation_datanepali_to_english_pipeline_evaluation
Nepali-English Speech-to-Text Translation Evaluation Dataset
Dataset Description
This dataset is designed for evaluating Nepali→English speech-to-text translation pipelines.
It contains audio recordings of 300 Nepali sentences, spoken by three speakers, covering a range of sentence types (statements, questions, commands, complex sentences, and named entities/numbers).
Each sentence is paired with:
Source text (Nepali) transcription
Reference English translation
Audio… See the full description on the dataset page: https://huggingface.co/datasets/iamTangsang/nepali_to_english_pipeline_evaluation.
