CoolFace
Datasetpublic

eloukas/edgar-corpus

The dataset contains annual filings (10K) of all publicly traded firms from 1993-2020. The table data is stripped but all text is retained. This dataset allows easy access to the EDGAR-CORPUS dataset based on the paper EDGAR-CORPUS: Billions of Tokens Make The World Go Round (See References in README.md for details).

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
63likes7.1kdownloads
Dataset Card

Dataset Card for [EDGAR-CORPUS]

Table of Contents

Dataset Description

  • Point of Contact: Lefteris Loukas

Dataset Summary

This dataset card is based on the paper EDGAR-CORPUS: Billions of Tokens Make The World Go Round authored by Lefteris Loukas et.al, as published in the ECONLP 2021 workshop.

This dataset contains the annual reports of public companies from 1993-2020 from SEC EDGAR filings.

There is supported functionality to load a specific year.

Care: since this is a corpus dataset, different train/val/test splits do not have any special meaning. It's the default HF card format to have train/val/test splits.

If you wish to load specific year(s) of specific companies, you probably want to use the open-source software which generated this dataset, EDGAR-CRAWLER: https://github.com/nlpaueb/edgar-crawler.

Citation

If this work helps or inspires you in any way, please consider citing the relevant paper published at the 3rd Economics and Natural Language Processing (ECONLP) workshop at EMNLP 2021 (Punta Cana, Dominican Republic):

@inproceedings{loukas-etal-2021-edgar,
    title = "{EDGAR}-{CORPUS}: Billions of Tokens Make The World Go Round",
    author = "Loukas, Lefteris  and
      Fergadiotis, Manos  and
      Androutsopoulos, Ion  and
      Malakasiotis, Prodromos",
    booktitle = "Proceedings of the Third Workshop on Economics and Natural Language Processing",
    month = nov,
    year = "2021",
    address = "Punta Cana, Dominican Republic",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.econlp-1.2",
    pages = "13--18",
}

Supported Tasks

This is a raw dataset/corpus for financial NLP. As such, there are no annotations or labels.

Languages

The EDGAR Filings are in English.

Dataset Structure

Data Instances

Refer to the dataset preview.

Data Fields

filename: Name of file on EDGAR from which the report was extracted.<br> cik: EDGAR identifier for a firm.<br> year: Year of report.<br> section_1: Corressponding section of the Annual Report.<br> section_1A: Corressponding section of the Annual Report.<br> section_1B: Corressponding section of the Annual Report.<br> section_2: Corressponding section of the Annual Report.<br> section_3: Corressponding section of the Annual Report.<br> section_4: Corressponding section of the Annual Report.<br> section_5: Corressponding section of the Annual Report.<br> section_6: Corressponding section of the Annual Report.<br> section_7: Corressponding section of the Annual Report.<br> section_7A: Corressponding section of the Annual Report.<br> section_8: Corressponding section of the Annual Report.<br> section_9: Corressponding section of the Annual Report.<br> section_9A: Corressponding section of the Annual Report.<br> section_9B: Corressponding section of the Annual Report.<br> section_10: Corressponding section of the Annual Report.<br> section_11: Corressponding section of the Annual Report.<br> section_12: Corressponding section of the Annual Report.<br> section_13: Corressponding section of the Annual Report.<br> section_14: Corressponding section of the Annual Report.<br> section_15: Corressponding section of the Annual Report.<br>

python
import datasets

# Load the entire dataset
raw_dataset = datasets.load_dataset("eloukas/edgar-corpus", "full")

# Load a specific year and split
year_1993_training_dataset = datasets.load_dataset("eloukas/edgar-corpus", "year_1993", split="train")

Data Splits

ConfigTrainingValidationTest
full176,28922,05022,036
year_19931,060133133
year_19942,083261260
year_19954,110514514
year_19967,589949949
year_19978,0841,0111,011
year_19988,0401,0061,005
year_19997,864984983
year_20007,589949949
year_20017,181898898
year_20026,636830829
year_20036,672834834
year_20047,111889889
year_20057,113890889
year_20067,064883883
year_20076,683836835
year_20087,408927926
year_20097,336917917
year_20107,013877877
year_20116,724841840
year_20126,479810810
year_20136,372797796
year_20146,261783783
year_20156,028754753
year_20165,812727727
year_20175,635705704
year_20185,508689688
year_20195,354670669
year_20205,480686685

Dataset Creation

Source Data

Initial Data Collection and Normalization

Initial data was collected and processed by the authors of the research paper EDGAR-CORPUS: Billions of Tokens Make The World Go Round.

Who are the source language producers?

Public firms filing with the SEC.

Annotations

Annotation process

NA

Who are the annotators?

NA

Personal and Sensitive Information

The dataset contains public filings data from SEC.

Considerations for Using the Data

Social Impact of Dataset

Low to none.

Discussion of Biases

The dataset is about financial information of public companies and as such the tone and style of text is in line with financial literature.

Other Known Limitations

The dataset needs further cleaning for improved performance.

Additional Information

Licensing Information

EDGAR data is publicly available.

Shoutout

Huge shoutout to @JanosAudran for the HF Card setup!

References

  • [Research Paper] Lefteris Loukas, Manos Fergadiotis, Ion Androutsopoulos, and, Prodromos Malakasiotis. EDGAR-CORPUS: Billions of Tokens Make The World Go Round. Third Workshop on Economics and Natural Language Processing (ECONLP). https://arxiv.org/abs/2109.14394 - Punta Cana, Dominican Republic, November 2021.
  • [Software] Lefteris Loukas, Manos Fergadiotis, Ion Androutsopoulos, and, Prodromos Malakasiotis. EDGAR-CRAWLER. https://github.com/nlpaueb/edgar-crawler (2021)
  • [EDGAR CORPUS, but in zip files] EDGAR CORPUS: A corpus for financial NLP research, built from SEC's EDGAR. https://zenodo.org/record/5528490 (2021)
  • [Word Embeddings] EDGAR-W2V: Word2vec Embeddings trained on EDGAR-CORPUS. https://zenodo.org/record/5524358 (2021)
  • [Applied Research paper where EDGAR-CORPUS is used] Lefteris Loukas, Manos Fergadiotis, Ilias Chalkidis, Eirini Spyropoulou, Prodromos Malakasiotis, Ion Androutsopoulos, and, George Paliouras. FiNER: Financial Numeric Entity Recognition for XBRL Tagging. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). https://doi.org/10.18653/v1/2022.acl-long.303 (2022)