CoolFace
Datasetpublic

BrunoHays/english-x-code-switching-samples

Synthetic English Code-Switching Evaluation Set Samples This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes43downloads
Dataset Card

Synthetic English Code-Switching Evaluation Set Samples

This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset.

Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row.

Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99. The samples dataset stores those same normalized chunks with parent_id links back to the mixed sample.