CoolFace
Modelpublic

ginic/vary_individuals_young_only_3_wav2vec2-large-xlsr-53-buckeye-ipa

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes14downloads
Model Card

license: mit language:

  • —en pipeline_tag: automatic-speech-recognition ---

About

This model was created to support experiments for evaluating phonetic transcription with the Buckeye corpus as part of https://github.com/ginic/multipa. This is a version of facebook/wav2vec2-large-xlsr-53 fine tuned on a specific subset of the Buckeye corpus. For details about specific model parameters, please view the config.json here or training scripts in the scripts/buckeye_experiments folder of the GitHub repository.

Experiment Details

These experiments keep the total amount of data equal to half the training data with the gender split 50/50, but further exclude certain speakers completely using the --speaker_restriction argument. This allows us to restrict speakers included in training data in any way. For the purposes of these experiments, we are focussed on the age demogrpahic of the user.

For reference, the speakers and their demographics included in the training data are as follows where the speaker age range 'y' means under 30 and 'o' means over 40:

speaker_idspeaker_genderspeaker_age_range
S01fy
S04fy
S08fy
S09fy
S12fy
S21fy
S02fo
S05fo
S07fo
S14fo
S16fo
S17fo
S06my
S11my
S13my
S15my
S28my
S30my
S03mo
S10mo
S19mo
S22mo
S24mo

Goals:

  • —Determine how variety of speakers in the training data affects performance

Params to vary:

  • —training seed (--train_seed)
  • —demographic make up of training data by age, using --speaker_restriction
  • —Experiments young_only: only individuals under 30, S01 S04 S08 S09 S12 S21 S06 S11 S13 S15 S28 S30
  • —Experiments old_only: only individuals over 40, S02 S05 S07 S14 S16 S17 S03 S10 S19 S22 S24