datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dartdoc
Dartdoc - 한국 금융공시 텍스트 데이터셋
한국 금융감독원 전자공시시스템(DART) OpenAPI를 통해 수집한 한국어 LLM 학습용 데이터셋입니다.
사업보고서, 증권신고서 등 공시 문서에서 고품질 한국어 텍스트를 추출하였습니다.
데이터셋 개요
항목
내용
언어
한국어 (ko)
수집 기간
2020년 ~ 2025년
총 레코드 수
256,548건
총 텍스트
약 4.6억 자
평균 청크 길이
약 1,794자
출처
금융감독원 DART OpenAPI
수집 대상
공시 유형
코드
대상 문서
필터 조건
정기공시
A
사업보고서
반기/분기보고서 제외
발행공시
C
증권신고서
정정신고서·집합투자 제외
추출 섹션
문서 전체가 아닌 품질이 높은 본문 섹션만 추출합니다.
섹션
내용
II
사업의 내용
IV
이사의 경영진단 및… See the full description on the dataset page: https://huggingface.co/datasets/chaannwooff/Dartdoc.adaption-preference-trace-decisions
PreferenceTrace — Source Corpus and Adaption Export
PreferenceTrace tests exact decision-making under competing preferences, evidence, approvals, abstention requirements, temporal/contextual precedence, and machine-readable citation contracts.
Two explicit lineage artifacts
File
Rows
Role
SHA-256
preferencetrace-source-96.jsonl
96
Canonical PreferenceTrace source corpus
7a447f9bf47c3ea455ed96ec36860360aa0e7b9e2dc604450e3a1c665b52363e… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-preference-trace-decisions.
