datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EDGAR_FILINGS_DATASET
SFD: SEC Filings Dataset (v1)
SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation.
This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in:
The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.sec-filings-qa-instruct
SEC Filings Instruction-Tuning Dataset (Llama-3 Format)
This dataset contains 5,000 curated, instruction-formatted question-answering pairs derived from corporate SEC filings (Forms 10-K and 10-Q). It is structured specifically for parameter-efficient instruction fine-tuning (SFT/QLoRA) of Small Language Models using the standard Llama-3 ChatML template.
Dataset Details
Origin Source: Curated subset extracted from nvidia/Nemotron-SpecializedDomains-Finance-v1.… See the full description on the dataset page: https://huggingface.co/datasets/lateesha-bhatia/sec-filings-qa-instruct.cakradana-kpu-filings-14k-pages
🏛️ Cakradana — KPU Campaign-Finance Filings
Indonesian election campaign-finance filings, for document OCR and structured extraction
📋 Table of Contents
🎯 What this is
📦 What is in here
🚀 Loading
🗂️ Schema
🔒 Redaction
⚠️ Limitations
📜 Licence and provenance
🎯 What this is
Indonesian candidates and parties must file campaign-finance reports — LADK,
LPSDK and LPPDK — with the KPU (Komisi Pemilihan Umum, the General
Elections… See the full description on the dataset page: https://huggingface.co/datasets/cakradana-app/cakradana-kpu-filings-14k-pages.
