datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system.
Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences.
Sentiment labels are provided on a per filing basis from the market reaction around the filing data.
Additional metadata for each filing is included in the dataset.financial-reports
Description
Topic: Financial Reports
Domains: Finance, Accounting, Economics
Focus: Synthetic raw financial reports for analysis and training
Number of Entries: 1000
Dataset Type: Raw Dataset
Model Used: bedrock/us.amazon.nova-pro-v1:0
Language: English
Generated by: SynthGenAI Package
vnpdf-financial-reports-dataset
Dataset Card for Financial Report Dataset Demo
Dataset Summary
This dataset contains financial reports from Vietnamese VININDEX, including both the text content and corresponding page images. The dataset is designed for document understanding and information extraction tasks.
Languages
The dataset contains text in Vietnamese (vi).
Dataset Structure
The dataset contains 401 examples, each with:
image: The page image from the PDF document
text: The… See the full description on the dataset page: https://huggingface.co/datasets/kiethuynhanh/vnpdf-financial-reports-dataset.financial-reports-extractive-summarization_eval
Financial Reports Extractive Summarization Evaluation Dataset
Validation and test splits for evaluating models on Arabic financial reports extractive summarization.
Dataset Structure
Format: Simple prompt-answer pairs
Validation: ~20 examples (10%)
Test: ~20 examples (10%)
Language: Arabic
Domain: Financial reports and market news
Fields
id: Unique identifier
prompt: The summarization prompt
full_text: Complete financial report
answer: Ground… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/financial-reports-extractive-summarization_eval.Financial-Reportsfrench_financial_reports_unfilteredfinancial_reportsfinancial-reports-extractive-summarization_train
Financial Reports Extractive Summarization Training Dataset
Training split of the Arabic financial reports extractive summarization dataset in conversational format.
Dataset Structure
Format: Conversational (human-agent pairs)
Size: ~160 training examples (80% of total)
Language: Arabic
Domain: Financial reports and market news
Features
id: Unique identifier
conversations: Human prompt and agent summary
report_type: Type of financial report… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/financial-reports-extractive-summarization_train.Vision-OCR-Financial-Reports-10kfinancial_reports_jsonfinancial_reports_csv
