chaannwooff/Dartdoc
Dartdoc - 한국 금융공시 텍스트 데이터셋 한국 금융감독원 전자공시시스템(DART) OpenAPI를 통해 수집한 한국어 LLM 학습용 데이터셋입니다. 사업보고서, 증권신고서 등 공시 문서에서 고품질 한국어 텍스트를 추출하였습니다. 데이터셋 개요 항목 내용 언어 한국어 (ko) 수집 기간 2020년 ~ 2025년 총 레코드 수 256,548건 총 텍스트 약 4.6억 자 평균 청크 길이 약 1,794자 출처 금융감독원 DART OpenAPI 수집 대상 공시 유형 코드 대상 문서 필터 조건 정기공시 A 사업보고서 반기/분기보고서 제외 발행공시 C 증권신고서 정정신고서·집합투자 제외 추출 섹션 문서 전체가 아닌 품질이 높은 본문 섹션만 추출합니다. 섹션 내용 II 사업의 내용… See the full description on the dataset page: https://huggingface.co/datasets/chaannwooff/Dartdoc.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face