CoolFace
Modelpublic

MahmoodAnaam/flaird-modernbert-large-attention-single-task

sourceHugging Faceupdated 5d agoView on Hugging Face
0likes112downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

flaird-modernbert-large-attention-single-task

This model is a fine-tuned version of [](https://huggingface.co/) on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 0.0018
  • Roc-auc: 0.999
  • Brier: 0.991
  • C@1: 0.987
  • F1: 0.993
  • F05u: 0.997
  • Mean: 0.994

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • trainbatchsize: 128
  • evalbatchsize: 128
  • seed: 42
  • optimizer: Use OptimizerNames.ADAMWTORCHFUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lrschedulertype: cosine
  • lrschedulerwarmup_steps: 0.06
  • num_epochs: 1

Training results

Training LossEpochStepValidation LossRoc-aucBrierC@1F1F05uMean
0.00700.100039680.00670.9930.9670.9560.9770.990.977
0.00530.200079360.00490.9960.9640.9490.9730.9890.974
0.00390.3001119040.00420.9980.9850.9790.9890.9950.989
0.00330.4001158720.00410.9980.9660.9540.9760.990.977
0.00280.5001198400.00290.9990.9850.9790.9890.9950.99
0.00220.6001238080.00260.9990.9830.9770.9880.9950.988
0.00210.7001277760.00210.9990.990.9860.9930.9970.993
0.00180.8002317440.00190.9990.9890.9850.9920.9970.992
0.00150.9002357120.00180.9990.9910.9870.9930.9970.994
0.00161.0396720.00180.9990.9910.9870.9930.9970.994

Framework versions

  • Transformers 5.16.1
  • Pytorch 2.11.0+cu128
  • Datasets 4.0.0
  • Tokenizers 0.23.1