datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KOREAN-WEBTEXT
KOREAN-WEBTEXT
KOREAN-WEBTEXT is a high-quality Korean language corpus consisting of 2.2 billion tokens. The data has been collected from the following sources:
cc100
oscar-corpus/OSCAR-2201
oscar-corpus/OSCAR-2109
oscar-corpus/OSCAR-2301
ontocord/CulturaY
Additional credible internet sources collected by out team
(We are working to add more sources)
The dataset undergoes rigorous filtering at both the sentence and document levels to ensure quality of text data. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KOREAN-WEBTEXT.cosmopedia_web_textbooksko-parallel-webtextcosmopedia_web_textbooks_logprobstask1728_web_nlg_data_to_text
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.webtextqa_zhFrom 2015 to 2016
korean-webtext-edu
🇰🇷🌐📚 korean-webtext-edu
HAERAE-HUB/KOREAN-WEBTEXT를 devngho/ko_edu_classifier_v2_nlpai-lab_KoE5 모델로 평가한 데이터셋
불러오기
from datasets import load_dataset
ds = load_dataset("devngho/korean-webtext-edu", name="scored_over_3", split="train")
성능
예정
컴퓨팅
Google Cloud TPU, transformers, JAX, tpuswarm
하드웨어
TPU v4-8 x 4 instances, 약 35분 소요
이 연구는 Google의 TPU Research Cloud (TRC)의 Cloud TPU 제공으로 수행되었습니다. ⚡
라이선스
원본… See the full description on the dataset page: https://huggingface.co/datasets/devngho/korean-webtext-edu.luau-org-web-textKOREAN-WEBTEXTopen_web_text_synthetic_queriestiny-webtext
Tiny WebText
The Tiny WebText dataset is designed to help models learn about perception on web text while neutralizing the bias of the source text using critical thinking methods. By providing a rich and diverse set of texts, I aim to improve the ability of models to understand and analyze information in a more objective and unbiased manner.
This dataset can be used to train and evaluate natural language processing and machine learning models, with the goal of improving their… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-webtext.spoken-web-questions-textgerman-webtext-quality-classification-dataset
Dataset Card for Dataset Name
Train and (manually annotated) test data of Paper:
Bootstrapping a Sentence-Level Corpus Quality Classifier for Web Text using Active Learning (RANLP25)
Dataset Details
see: https://aclanthology.org/2025.globalnlp-1.12/
open-web-textkorean_unlabeled_web_textweb_text_synthetic_dataset_50kflan_combined_task1728_web_nlg_data_to_textwebtextkorean_webtext_10000text_L2-regular-TTS_spoken-web-questionstext_merge-dare_spoken-web-questionsspoken-web-questions-text-scoretext_L2-regular-15_spoken-web-questionstext_L2-regular-ASR_spoken-web-questions-scoretext_L2-regular-linear_spoken-web-questionsmini-webtextspoken-web-questions-text_original-scoretext_llama-origin_spoken-web-questionswebtext_dates_sortedspoken-web-questions-text_original
