CoolFace
Datasetpublic

chaannwooff/Dartdoc

Dartdoc - 한국 금융공시 텍스트 데이터셋 한국 금융감독원 전자공시시스템(DART) OpenAPI를 통해 수집한 한국어 LLM 학습용 데이터셋입니다. 사업보고서, 증권신고서 등 공시 문서에서 고품질 한국어 텍스트를 추출하였습니다. 데이터셋 개요 항목 내용 언어 한국어 (ko) 수집 기간 2020년 ~ 2025년 총 레코드 수 256,548건 총 텍스트 약 4.6억 자 평균 청크 길이 약 1,794자 출처 금융감독원 DART OpenAPI 수집 대상 공시 유형 코드 대상 문서 필터 조건 정기공시 A 사업보고서 반기/분기보고서 제외 발행공시 C 증권신고서 정정신고서·집합투자 제외 추출 섹션 문서 전체가 아닌 품질이 높은 본문 섹션만 추출합니다. 섹션 내용 II 사업의 내용… See the full description on the dataset page: https://huggingface.co/datasets/chaannwooff/Dartdoc.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
1likes152downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
chaannwooff/Dartdoc · CoolFace