CoolFace
Datasetpublic

BrunoHays/english-x-code-switching

Synthetic English Code-Switching Evaluation Set This dataset contains synthetic long-form English code-switching audio samples built from ML-SUPERB hybrid data. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching.

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes113downloads
Dataset Card

Synthetic English Code-Switching Evaluation Set

This dataset contains synthetic long-form English code-switching audio samples built from ML-SUPERB hybrid data.

Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row.

Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99. The samples dataset stores those same normalized chunks with parent_id links back to the mixed sample.