web-text
Datasets
All datasets matching “web-text”KOREAN-WEBTEXT
KOREAN-WEBTEXT
KOREAN-WEBTEXT is a high-quality Korean language corpus consisting of 2.2 billion tokens. The data has been collected from the following sources:
cc100
oscar-corpus/OSCAR-2201
oscar-corpus/OSCAR-2109
oscar-corpus/OSCAR-2301
ontocord/CulturaY
Additional credible internet sources collected by out team
(We are working to add more sources)
The dataset undergoes rigorous filtering at both the sentence and document levels to ensure quality of text data. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KOREAN-WEBTEXT.cosmopedia_web_textbookskorean-webtext-edu
Korean Webtext Edu
데이터셋 설명
Korean-webtext-edu는 HAERAE-HUB/KOREAN-WEBTEXT 데이터셋에서 선별된 하위 집합입니다. 이 데이터셋은 원본 데이터의 128만 개 문서 중 교육적 가치가 높은 콘텐츠만을 필터링하여 구축되었습니다.
본 데이터셋의 필터링 과정은 HuggingFaceFW/fineweb-edu 데이터셋의 방법론에서 영감을 받았습니다. 교육적이고 사실적이며 구조화된 한국어 웹 텍스트를 대규모로 제공하여, 모델 학습의 질을 높이는 것을 목표로 합니다.
데이터셋 구축
소스 데이터
HAERAE-HUB/KOREAN-WEBTEXT
필터링 및 점수 산정
"교육적 가치" 점수 산정 방식은 FineWeb-edu 방법론을 기반으로 합니다. 텍스트의 일관성, 사실 관계의 정확성, 핵심 개념 소개, 교육적 적합성 등을… See the full description on the dataset page: https://huggingface.co/datasets/eliceai/korean-webtext-edu.cosmopedia_web_textbooks_logprobsko-parallel-webtexttask1728_web_nlg_data_to_text
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.
