datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Numb3rs
Numb3rs - Numbers Speech Benchmark (Dataset)
A speech dataset for text normalization (TN) and inverse text normalization (ITN) tasks, containing paired written/spoken forms with corresponding synthetic audio.
Dataset Creation
This dataset was created through the following pipeline:
Source Data: Text normalization pairs were derived from the Google Text Normalization dataset, containing written forms (e.g., "$100") and their spoken equivalents (e.g., "one hundred… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Numb3rs.nvidia-brain-noise-evaluation-dataset
Nvidia Brain Noise Evaluation Dataset
Dataset Description
This dataset contains 64 samples organized across multiple splits and 32 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
noisy-bg-snr-10: 2 samples
test: 2 samples
noisy-bg-snr-20: 2 samples
test: 2 samples
noisy-bg-snr-30: 2 samples
test: 2 samples
noisy-bg-snr-40: 2 samples
test: 2 samples
noisy-bg-snr-50: 2 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/nvidia-brain-noise-evaluation-dataset.
