CoolFace
Datasetpublic

nglaura/pubmedlay-summarization

LoRaLay: A Multilingual and Multimodal Dataset for Long Range and Layout-Aware Summarization A collaboration between reciTAL, MLIA (ISIR, Sorbonne Université), Meta AI, and Università di Trento PubMed-Lay dataset for summarization PubMed-Lay is an enhanced version of the PubMed summarization dataset, for which layout information is provided. Data Fields article_id: article id article_words: sequence of words constituting the body of the article… See the full description on the dataset page: https://huggingface.co/datasets/nglaura/pubmedlay-summarization.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
0likes229downloads
Dataset Card

LoRaLay: A Multilingual and Multimodal Dataset for Long Range and Layout-Aware Summarization

A collaboration between reciTAL, MLIA (ISIR, Sorbonne Université), Meta AI, and Università di Trento

PubMed-Lay dataset for summarization

PubMed-Lay is an enhanced version of the PubMed summarization dataset, for which layout information is provided.

Data Fields

  • article_id: article id
  • article_words: sequence of words constituting the body of the article
  • article_bboxes: sequence of corresponding word bounding boxes
  • norm_article_bboxes: sequence of corresponding normalized word bounding boxes
  • abstract: a string containing the abstract of the article
  • article_pdf_url: URL of the article's PDF

Data Splits

This dataset has 3 splits: train, validation, and test.

Dataset SplitNumber of Instances
Train78,234
Validation4,084
Test4,350

Citation

latex
@article{nguyen2023loralay,
  title={LoRaLay: A Multilingual and Multimodal Dataset for Long Range and Layout-Aware Summarization},
  author={Nguyen, Laura and Scialom, Thomas and Piwowarski, Benjamin and Staiano, Jacopo},
  journal={arXiv preprint arXiv:2301.11312},
  year={2023}
}