CoolFace
Modelpublic

wandb/sourcecode-detection

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes14downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

CodeBERTa-small-v1-sourcecode-detection-clf

This model is a fine-tuned version of huggingface/CodeBERTa-small-v1 on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 0.0171
  • F1: 0.9975
  • Accuracy: 0.9975
  • Precision: 0.9975
  • Recall: 0.9975

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0003
  • trainbatchsize: 320
  • evalbatchsize: 320
  • seed: 2024
  • optimizer: Use OptimizerNames.ADAMWTORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizerargs=No additional optimizer arguments
  • lrschedulertype: cosine
  • lrschedulerwarmup_ratio: 0.1
  • num_epochs: 1

Training results

Training LossEpochStepValidation LossF1AccuracyPrecisionRecall
No log000.69810.33370.50010.61620.5001
0.02940.142010000.03980.99470.99470.99470.9947
0.00760.284120000.02110.99680.99680.99680.9968
0.00530.426130000.01880.99730.99730.99730.9973
0.00560.568140000.01660.99760.99760.99760.9976
0.00440.710150000.01720.99750.99750.99750.9975
0.00090.852260000.01710.99750.99750.99750.9975
0.00520.994270000.01710.99750.99750.99750.9975

Framework versions

  • Transformers 4.46.3
  • Pytorch 2.5.1
  • Datasets 3.1.0
  • Tokenizers 0.20.3