CoolFace
Datasetpublic

ultrasafe-ai/asr-eval-noise-clean-mix-v1-1k

ASR Eval — Restaurant Speech v1 (1K) A 1,000-sample English speech benchmark dataset recorded in a real restaurant environment, designed to evaluate ASR systems under challenging real-world acoustic conditions. Dataset Summary Each audio sample was recorded in a restaurant setting, capturing natural speech alongside the ambient sounds typical of a busy dining environment — background conversations, cutlery, and general crowd noise. This makes it an ideal benchmark… See the full description on the dataset page: https://huggingface.co/datasets/ultrasafe-ai/asr-eval-noise-clean-mix-v1-1k.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
1likes12downloads
Dataset Card

ASR Eval — Restaurant Speech v1 (1K)

A 1,000-sample English speech benchmark dataset recorded in a real restaurant environment, designed to evaluate ASR systems under challenging real-world acoustic conditions.

Dataset Summary

Each audio sample was recorded in a restaurant setting, capturing natural speech alongside the ambient sounds typical of a busy dining environment — background conversations, cutlery, and general crowd noise. This makes it an ideal benchmark for evaluating ASR robustness in real-world noisy environments.

Dataset Structure

ColumnTypeDescription
uuidstringUnique sample identifier
audioAudio (16 kHz)Speech recorded in a restaurant environment
durationfloatAudio duration in seconds
transcriptstringGround truth transcript

Usage

python
from datasets import load_dataset

ds = load_dataset("ultrasafe-ai/asr-eval-noise-clean-mix-v1-1k", split="test")

sample = ds[0]
print(sample["transcript"])
# → "this election is very important to..."

audio_array  = sample["audio"]["array"]         # numpy array, float32
sample_rate  = sample["audio"]["sampling_rate"] # 16000
duration_sec = sample["duration"]               # e.g. 9.32

Benchmarking Example

python
# Transcribe and compute WER
from jiwer import wer

reference  = sample["transcript"]
hypothesis = your_asr_model(sample["audio"]["array"])
score = wer(reference, hypothesis)
print(f"WER: {score:.3f}")

Dataset Stats

  • —Samples: 1,000
  • —Language: English
  • —Sample rate: 16,000 Hz (mono)
  • —Avg duration: ~8–12 seconds per sample
  • —Recording environment: Restaurant (real-world ambient noise)