CoolFace
Modelpublic

esuriddick/led-base-16384-finetuned-govreport

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
0likes168downloads
Model Card

led-base-16384-finetuned-govreport

This model is a fine-tuned version of allenai/led-base-16384 on the pszemraj/govreport-summarization-8192 dataset. It achieves the following results on the evaluation set:

  • —Loss: 1.2887

The rouge metrics calculations were processed later down the line (final notebook can be found HERE).

It achieved the following results on the validation set:

  • —Rouge1: 50.3574
  • —Rouge2: 20.0448
  • —Rougel: 22.2156
  • —Rougelsum: 22.2156

It achieved the following results on the test set:

  • —Rouge1: 52.6378
  • —Rouge2: 22.2130
  • —Rougel: 23.5898
  • —Rougelsum: 23.5898

Model description

As described in Longformer: The Long-Document Transformer by Iz Beltagy, Matthew E. Peters, Arman Cohan, Allenai's Longformer Encoder-Decoder (LED) was initialized from *bart-base* since both models share the exact same architecture. To be able to process 16K tokens, bart-base's position embedding matrix was simply copied 16 times.

This model is especially interesting for long-range summarization and question answering.

Intended uses & limitations

pszemraj/govreport-summarization-8192 is a pre-processed version of the dataset ccdv/govreport-summarization, which is a dataset for summarization of long documents adapted from this repository and this paper.

The Allenai's LED model was fine-tuned to this dataset, allowing the summarization of documents up to 16384 tokens.

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 5e-05
  • —trainbatchsize: 1
  • —evalbatchsize: 1
  • —seed: 42
  • —gradientaccumulationsteps: 8
  • —totaltrainbatch_size: 8
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: linear
  • —num_epochs: 2

Training results

Training LossEpochStepValidation Loss
1.14920.242501.4233
1.00770.495001.3813
1.00690.737501.3499
0.96390.9810001.3216
0.79961.2212501.3172
0.93951.4615001.3003
0.9131.7117501.2919
0.88431.9520001.2887

Framework versions

  • —Transformers 4.30.2
  • —Pytorch 2.0.0
  • —Datasets 2.1.0
  • —Tokenizers 0.13.3