thiagolira/CiceroASR
<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->
CiceroASR
This model is a fine-tuned version of facebook/w2v-bert-2.0 for the transcription of Classical Latin!
Example from the Aeneid: <video controls src="https://cdn-uploads.huggingface.co/production/uploads/5fc7944e8a82cc0bcf7cc51d/hYNFr2od1EKDlRRdzJmzR.webm"></video> Transcription: arma virumque cano (Of arms and men I sing)
Example from Genesis: <video controls src="https://cdn-uploads.huggingface.co/production/uploads/5fc7944e8a82cc0bcf7cc51d/9Q6DfG2h8FkABnl55DLBH.webm"></video> Transcription (little error there): creavit deus chaelum et terram (In the beggining God created the heaven and the earth)
It achieves the following results on the evaluation set of my dataset Latin Youtube:
- Loss: 0.5395
- Wer: 0.2220
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 0.0001
- trainbatchsize: 16
- evalbatchsize: 8
- seed: 42
- gradientaccumulationsteps: 2
- totaltrainbatch_size: 32
- optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
- lrschedulertype: linear
- lrschedulerwarmup_steps: 300
- num_epochs: 15
- mixedprecisiontraining: Native AMP
Training results
Framework versions
- Transformers 4.38.1
- Pytorch 2.1.0+cu121
- Datasets 2.17.1
- Tokenizers 0.15.2
