CoolFace
Datasetpublic

sfd-anonymous/sec-parser

SEC Filings Dataset Parser This repository contains the core parser used to convert SEC EDGAR filings into layout-faithful Markdown-style text for downstream dataset construction and evaluation. Contents sec_parser/sec_parser.py: main parser implementation sec_parser/special_chars.py: special-character normalization tables sec_parser/hardcodes.py: filing cleanup hardcodes sec_parser/config.py: parser configuration pdf_table_fastpath.py, table_ocr_backends.py… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/sec-parser.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes89downloads
Dataset Card

SEC Filings Dataset Parser

This repository contains the core parser used to convert SEC EDGAR filings into layout-faithful Markdown-style text for downstream dataset construction and evaluation.

Contents

  • sec_parser/sec_parser.py: main parser implementation
  • sec_parser/special_chars.py: special-character normalization tables
  • sec_parser/hardcodes.py: filing cleanup hardcodes
  • sec_parser/config.py: parser configuration
  • pdf_table_fastpath.py, table_ocr_backends.py, mistral_pdf_ocr_overlay.py: PDF/table OCR helper code
  • test_sec_parser.py: smoke/regression tests
  • Dockerfile, Dockerfile.linux, requirements.txt: reproducible environment files

Quick Start

bash
pip install -r requirements.txt
python -m unittest test_sec_parser.py
python sec_parser/sec_parser.py path/to/filing.txt

Some PDF/OCR paths require external OCR credentials and browser rendering dependencies.