datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisperkit-evals
WhisperKit
WhisperKit is an on-device speech recognition framework for Apple Silicon:
https://github.com/argmaxinc/WhisperKit
For performance and accuracy benchmarks on real devices, please see:
https://huggingface.co/spaces/argmaxinc/whisperkit-benchmarks
whisperkit-evals-dataset
WhisperKit Evals Dataset
Overview
The WhisperKit Evals Dataset is a comprehensive collection of our speech recognition evaluation results, specifically designed to benchmark the performance of WhisperKit models across various devices and operating systems. This dataset provides detailed insights into performance and quality metrics, and model behavior under different conditions.
Dataset Structure
The dataset is organized into JSON files, each representing a… See the full description on the dataset page: https://huggingface.co/datasets/argmaxinc/whisperkit-evals-dataset.whisperkit-test-datawhisperkit-evals-multilingual
WhisperKit Evaluation Results
Dataset: common_voice_17_0-argmax_subset-400
Short-form Audio (<30s/clip) - Max 400 samples per language from Common Voice 17.0 Test Set
es
ro
th
nl
id
sv
de
pl
fi
it
cs
en
vi
el
hu
ru
gl
fr
pt
da
File Size (MB)
Code Commit
WhisperKit/openai_whisper-large-v3
4.93
5.39
6.11
7.03
9.47
9.81
9.89
10.13
10.32
11.11… See the full description on the dataset page: https://huggingface.co/datasets/argmaxinc/whisperkit-evals-multilingual.whisperkit-evals_01-30-24
WhisperKit Evaluation Results
Dataset: librispeech
WhisperKit + openai_whisper-large-v3 (+optimized variants)
WER
QoI (%)
File Size (MB)
openai_whisper-large-v3
2.44
100
3100
openai_whisper-large-v3_turbo
2.41
99.8
3100
openai_whisper-large-v3_turbo_1307MB
2.6
97.7
1307
openai_whisper-large-v3_turbo_1049MB
4.81
91
1049
openai_whisper-large-v3_1053MB
4.65
90.8
1053
Different Projects + openai_whisper-large-v3
WER… See the full description on the dataset page: https://huggingface.co/datasets/argmaxinc/whisperkit-evals_01-30-24.whisperkit-0.7.0-evals
WhisperKit-0.7.0 VAD Chunking Strategy Evaluation Results
This is an evaluation study to verify that the Voice Activity Detection (VAD) based chunk-and-batch strategy introduced in WhisperKit-0.7.0 does not decrease transcription quality. In order to measure the impact of chunking, we picked a random 10% subset of the earnings22 dataset which comprises corporate earnings call recordings in English with various accents. The long-form nature (>1hr/clip) and the density of speech in… See the full description on the dataset page: https://huggingface.co/datasets/argmaxinc/whisperkit-0.7.0-evals.whisperkit_testsAll files are from: earnings22
Rencoded to 24kbps MP3 using:
ffmpeg -i 4446796.wav -vn -map_metadata -1 -ac 1 -c:a libmp3lame -b:a 24k -application voip -y 4446796.mp3
