chaannwooff/CC-korean
CC-korean 한국어 웹 텍스트 코퍼스. CommonCrawl WET 파일에서 추출한 한국어 문서 데이터셋입니다. 데이터 수집 방법 소스: CommonCrawl WET 파일 (HTML 제거된 텍스트 포맷) 수집 스냅샷: CC-MAIN-2025-05 ~ CC-MAIN-2026-12 (총 15개 스냅샷, 2025~2026년) 언어 필터: fasttext lid.176.bin 모델로 한국어(ko) 감지, 신뢰도 0.5 이상만 수집 1차 품질 필터: 수집 단계에서 휴리스틱 줄 단위 노이즈 제거 + KenLM perplexity 필터 적용 필터링 수집 후 추가로 2단계 필터링 적용. 줄 단위 인라인 제거: URL, 이모티콘, 대괄호 태그, 기자 바이라인, 출처 표기 등 줄 단위 제거 항목: 한국어 없는 줄 (영문·중문·일문 등) UI/네비게이션 잔재 (로그인, 더보기, 공유 버튼, 파이프 구분자… See the full description on the dataset page: https://huggingface.co/datasets/chaannwooff/CC-korean.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face