CoolFace
Datasetpublic

Luigi/breezyvoice-zhtw-en-phone-corpus

BreezyVoice zh-TW / English Phone-Attendant Corpus (ASR-verified) Synthetic Taiwan-Mandarin + English code-mixed speech for a phone-attendant domain, generated with MediaTek-Research/BreezyVoice (zero-shot voice clone, fixed reference voice) and ASR-verified: every clip was transcribed with faster-whisper and kept only if its Han-character CER vs the intended text was below 0.3. Built to distill BreezyVoice into a tiny real-time on-device TTS (the Inflect-Nano architecture)… See the full description on the dataset page: https://huggingface.co/datasets/Luigi/breezyvoice-zhtw-en-phone-corpus.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes52downloads
Dataset Card

BreezyVoice zh-TW / English Phone-Attendant Corpus (ASR-verified)

Synthetic Taiwan-Mandarin + English code-mixed speech for a phone-attendant domain, generated with [MediaTek-Research/BreezyVoice](https://huggingface.co/MediaTek-Research/BreezyVoice) (zero-shot voice clone, fixed reference voice) and ASR-verified: every clip was transcribed with faster-whisper and kept only if its Han-character CER vs the intended text was below 0.3.

Built to distill BreezyVoice into a tiny real-time on-device TTS (the Inflect-Nano architecture), targeting 8 kHz G.711 phone audio on a Jetson Nano. Released so others can reuse the teacher data.

  • —9042 clips, 14.4 h, 22050 Hz mono
  • —Text: zh-TW / code-mix phone-attendant lines (transfers, names, order numbers, English terms)
  • —Each row: audio, text, duration, cer (ASR Han-CER vs text)
  • —Single synthetic speaker (BreezyVoice reference voice)

Provenance & license

Audio synthesized by BreezyVoice (Apache-2.0, MediaTek Research). This derivative corpus is released Apache-2.0. Please cite BreezyVoice (arXiv:2501.17790) and respect its license. Synthetic data — not human recordings.

Fields

fieldtypedescription
audioAudio22050 Hz mono wav
textstringintended text (zh-TW / code-mix)
durationfloatseconds
cerfloatfaster-whisper Han-CER vs text (lower = better match)