CoolFace
Datasetpublic

NLP-07-ODQA/kowiki-cleaned

NLP-07-ODQA/kowiki-cleaned Dataset Description 이 데이터셋은 한국어 위키피디아 XML 덤프에서 추출하고 정제한 데이터셋입니다. 청킹 전 단계의 정제된 문서를 포함합니다. 데이터 소스 원본: 한국어 위키피디아 XML 덤프 네임스페이스: 0 (일반 문서) 처리된 페이지 수: 565,484개 전체 페이지 수: 2,157,147개 데이터 정제 과정 위키 마크업 제거: [[링크]], {템플릿}, ==제목== 등 제거 불필요한 섹션 제거: '같이 보기', '외부 링크', '참고 문헌', '각주' 등 제거 필터링: 리다이렉트, 빈 페이지, 스텁 페이지 제거 최소 길이: 200자 이상만 포함 통계 평균 텍스트 길이: 1871.8자 데이터 구조 각 데이터 포인트는 다음 필드를 포함합니다:… See the full description on the dataset page: https://huggingface.co/datasets/NLP-07-ODQA/kowiki-cleaned.

sourceHugging Facecc-by-sa-4.0updated 9mo agoView on Hugging Face
0likes24downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
NLP-07-ODQA/kowiki-cleaned · CoolFace