CoolFace
Modelpublic

MahmoodAnaam/flaird-modernbert-large-attention-single-task-frozen

sourceHugging Faceupdated 10d agoView on Hugging Face
0likes119downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

flaird-modernbert-large-attention-single-task-frozen

This model is a fine-tuned version of [](https://huggingface.co/) on an unknown dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.0073
  • —Roc-auc: 0.991
  • —Brier: 0.956
  • —C@1: 0.936
  • —F1: 0.966
  • —F05u: 0.985
  • —Mean: 0.967

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 2e-05
  • —trainbatchsize: 256
  • —evalbatchsize: 256
  • —seed: 42
  • —optimizer: Use OptimizerNames.ADAMWTORCHFUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • —lrschedulertype: cosine
  • —lrschedulerwarmup_steps: 0.06
  • —num_epochs: 1

Training results

Training LossEpochStepValidation LossRoc-aucBrierC@1F1F05uMean
0.01610.100019840.01490.9640.9290.9120.9530.9780.947
0.01260.200039680.01130.9780.9320.8920.9410.9750.944
0.01110.300159520.01070.9820.9560.9410.9690.9860.967
0.01010.400179360.00960.9850.9310.8940.9420.9750.945
0.00920.500199200.00870.9870.9480.920.9570.9820.959
0.00870.6001119040.00960.9870.9260.8940.9420.9760.945
0.00820.7001138880.00780.990.9550.9350.9650.9850.966
0.00800.8002158720.00740.990.960.9430.970.9870.97
0.00770.9002178560.00730.9910.9590.9390.9680.9860.969
0.00771.0198360.00730.9910.9560.9360.9660.9850.967

Framework versions

  • —Transformers 5.16.1
  • —Pytorch 2.11.0+cu128
  • —Datasets 4.0.0
  • —Tokenizers 0.23.1