ultrasafe-ai/asr-eval-noise-clean-mix-v1-1k
ASR Eval — Restaurant Speech v1 (1K) A 1,000-sample English speech benchmark dataset recorded in a real restaurant environment, designed to evaluate ASR systems under challenging real-world acoustic conditions. Dataset Summary Each audio sample was recorded in a restaurant setting, capturing natural speech alongside the ambient sounds typical of a busy dining environment — background conversations, cutlery, and general crowd noise. This makes it an ideal benchmark… See the full description on the dataset page: https://huggingface.co/datasets/ultrasafe-ai/asr-eval-noise-clean-mix-v1-1k.
ASR Eval — Restaurant Speech v1 (1K)
A 1,000-sample English speech benchmark dataset recorded in a real restaurant environment, designed to evaluate ASR systems under challenging real-world acoustic conditions.
Dataset Summary
Each audio sample was recorded in a restaurant setting, capturing natural speech alongside the ambient sounds typical of a busy dining environment — background conversations, cutlery, and general crowd noise. This makes it an ideal benchmark for evaluating ASR robustness in real-world noisy environments.
Dataset Structure
Usage
from datasets import load_dataset
ds = load_dataset("ultrasafe-ai/asr-eval-noise-clean-mix-v1-1k", split="test")
sample = ds[0]
print(sample["transcript"])
# → "this election is very important to..."
audio_array = sample["audio"]["array"] # numpy array, float32
sample_rate = sample["audio"]["sampling_rate"] # 16000
duration_sec = sample["duration"] # e.g. 9.32Benchmarking Example
# Transcribe and compute WER
from jiwer import wer
reference = sample["transcript"]
hypothesis = your_asr_model(sample["audio"]["array"])
score = wer(reference, hypothesis)
print(f"WER: {score:.3f}")Dataset Stats
- Samples: 1,000
- Language: English
- Sample rate: 16,000 Hz (mono)
- Avg duration: ~8–12 seconds per sample
- Recording environment: Restaurant (real-world ambient noise)
