marcelovidigal/DeepSeek-R1-Distill-Qwen-1.5B-2-contract-sections-classification-v4-50
012
<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->
DeepSeek-R1-Distill-Qwen-1.5B-2-contract-sections-classification-v4-50
This model is a fine-tuned version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B on an unknown dataset. It achieves the following results on the evaluation set:
- Loss: 0.9474
- Accuracy Evaluate: 0.7023
- Precision Evaluate: 0.7283
- Recall Evaluate: 0.7032
- F1 Evaluate: 0.7003
- Accuracy Sklearn: 0.7023
- Precision Sklearn: 0.7252
- Recall Sklearn: 0.7023
- F1 Sklearn: 0.6989
- Acuracia Rotulo Objeto: 0.8554
- Acuracia Rotulo Obrigacoes: 0.7795
- Acuracia Rotulo Valor: 0.6017
- Acuracia Rotulo Vigencia: 0.5932
- Acuracia Rotulo Rescisao: 0.6676
- Acuracia Rotulo Foro: 0.9654
- Acuracia Rotulo Reajuste: 0.4306
- Acuracia Rotulo Fiscalizacao: 0.6435
- Acuracia Rotulo Publicacao: 0.8177
- Acuracia Rotulo Pagamento: 0.4457
- Acuracia Rotulo Casos Omissos: 0.8276
- Acuracia Rotulo Sancoes: 0.7339
- Acuracia Rotulo Dotacao Orcamentaria: 0.7802
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 1e-06
- trainbatchsize: 16
- evalbatchsize: 16
- seed: 42
- optimizer: Use OptimizerNames.ADAMWTORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizerargs=No additional optimizer arguments
- lrschedulertype: linear
- num_epochs: 50
- mixedprecisiontraining: Native AMP
Training results
Framework versions
- PEFT 0.14.0
- Transformers 4.48.3
- Pytorch 2.6.0+cu124
- Datasets 3.3.0
- Tokenizers 0.21.0
