CoolFace
Modelpublic

Thalesian/SciGPT-2-finetuned-papers

sourceHugging Faceapache-2.0updated 4y agoView on Hugging Face
0likes24downloads
Model Card

<!-- This model card has been generated automatically according to the information Keras had access to. You should probably proofread and complete it, then remove this comment. -->

Thalesian/SciGPT-2-finetuned-papers

This model is a fine-tuned version of distilgpt2 on an unknown dataset. It achieves the following results on the evaluation set:

  • —Train Loss: 1.8928
  • —Validation Loss: 2.0726
  • —Epoch: 19

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —optimizer: {'name': 'AdamWeightDecay', 'clipvalue': 0.7, 'learningrate': {'classname': 'ExponentialDecay', 'config': {'initiallearningrate': 0.0005, 'decaysteps': 500, 'decayrate': 0.95, 'staircase': False, 'name': None}}, 'decay': 0.0, 'beta1': 0.9, 'beta2': 0.999, 'epsilon': 1e-07, 'amsgrad': False, 'weightdecayrate': 0.1}
  • —training_precision: float32

Training results

Train LossValidation LossEpoch
1.89762.07290
1.89542.07281
1.89422.07252
1.89372.07263
1.89322.07274
1.89292.07275
1.89292.07276
1.89262.07267
1.89282.07268
1.89262.07269
1.89272.072610
1.89272.072611
1.89272.072612
1.89262.072613
1.89272.072614
1.89272.072615
1.89272.072616
1.89272.072617
1.89272.072618
1.89282.072619

Framework versions

  • —Transformers 4.25.1
  • —TensorFlow 2.8.0
  • —Datasets 2.8.0
  • —Tokenizers 0.13.2