CoolFace
Datasetpublic

Tobias-B/ipa_augmentation_cv11_training_segments_RM_BM_TM

These files list the relevant Common Voice 11 segments which were used to train the reference model and therefore also the baseline and target model in the selective augmentation approach. List of Training Segments for Selective Augmentation: https://huggingface.co/collections/Tobias-B/universal-phonetic-asr-models-selective-augmentation-680b5034c0729058fadcf1d6 These models were created to advance automatic phonetic transcription (APT) beyond the training transcription… See the full description on the dataset page: https://huggingface.co/datasets/Tobias-B/ipa_augmentation_cv11_training_segments_RM_BM_TM.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes7downloads
Dataset Card

These files list the relevant Common Voice 11 segments which were used to train the reference model and therefore also the baseline and target model in the selective augmentation approach.

List of Training Segments for Selective Augmentation:

https://huggingface.co/collections/Tobias-B/universal-phonetic-asr-models-selective-augmentation-680b5034c0729058fadcf1d6

These models were created to advance automatic phonetic transcription (APT) beyond the training transcription accuracy. The workflow to improve APT is called Selective Augmentation and was developed by Tobias Bystrich at Fraunhofer Institute IAIS and using resources of WestAI: Simulations were performed with computing resources granted by WestAI under project rwth1594.

The models in this project are the reference (RM), helper (HM), baseline (BM) and target model (TM) for the selective augmentation workflow. Additionally, for reimplementation, the provided list of training segments ensures that the RM can predict the highest quality reference transcriptions.

The RM closely corresponds to a reimplemented MultIPA model (https://github.com/ctaguchi/multipa).

The target model has greatly improved plosive phonation information when measured against the baseline model. This is achieved by augmenting the baseline training data with reliable phonation information from a Hindi helper model.

A technical overview can be found in the LREC 2026 paper by T. Bystrich, J. M. Pritzen, C. A. Schmidt, C. Wich-Reif: "Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping", also available as a pre-print on arXiv.

For more in-depth discussions about this approach and automatic phonetic transcriptions in general, you may consult Tobias Bystrich's master's thesis: "Multilingual Automatic Phonetic Transcription – a Linguistic Investigation of its Performance on German and Approaches to Improving the State of the Art". https://doi.org/10.24406/publica-4418