datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
c4-ko-cleaned-2이전 데이터셋에서 아쉬운 점이 많이 보여 조금 개선한 데이터셋 입니다.
원본 데이터셋: c4
파일 크기: 약 10gb
데이터 수: 2261464
color-filtered-c4
CoLoR-Filtered C4
This repo contains two datasets: color-filtered-c4-books and color-filtered-c4-down associated with the CoLoR-Filter paper.
Each dataset is a 64x filtered version of the C4 dataset from Raffel et al., 2019 that has been selected using the CoLoR-Filter algorithm for data selection.
Each dataset has about 2.7b tokens when using the allenai/eleuther-ai-gpt-neox-20b-pii-special tokenizer.
color-filtered-c4-books was selected to target books based on a small (25m token)… See the full description on the dataset page: https://huggingface.co/datasets/davidbrandfonbrener/color-filtered-c4.c4-ko-cleaned학교 점심시간 때 할 거 없어서 만든 c4를 정제한 데이터입니다. 다 하면 컴퓨터가 감당 못 할 거 같아서 전체 데이터의 1/10만 진행하였으며 아마 품질은 안 좋을 겁니다.
파일 크기: 약 3gb
데이터 수: 1847023
msm-v2-shared-c4-36k
MSM v2 shared C4 36k
Training-ready inputs for the second MSM run. Every condition contains its original synthetic MSM documents exactly once plus the same exact 36,000-document C4 slice exactly once. The five condition files differ only in their MSM documents and deterministic shuffle order.
Synthetic rows begin with <DOCTAG>\n and declare the same string in mask_prefix; C4 rows are untagged and declare an empty mask_prefix. The trainer must mask only the declared prefix tokens… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/msm-v2-shared-c4-36k.c4
C4
This repository hosts a copy of the widely used C4 dataset, a variant of the Colossal Clean Crawled Corpus designed for training and evaluating Large Language Models (LLMs) on news-like text.
C4 consists of cleaned web data from Common Crawl, specifically curated to contain more news-style content. This dataset is commonly used in language modeling tasks, text generation, and research focused on news and article-like content.
Contents
c4.json (or your actual… See the full description on the dataset page: https://huggingface.co/datasets/S3IC/c4.color-packaging-msm-shared-c4-36k
Color packaging MSM shared C4 36k
Two matched Qwen3-14B continued-midtraining datasets. Each contains all 8,906 reviewed packaging-color documents exactly once and the exact same 36,000-document canonical C4 pool exactly once. Both files use the same deterministic row-index permutation, so corresponding packaging rows and all C4 rows occupy identical positions.
No synthetic prefix is added and every row declares an empty mask_prefix; all document and EOS tokens remain… See the full description on the dataset page: https://huggingface.co/datasets/GaloisTheory123/color-packaging-msm-shared-c4-36k.c4-en-tokenized
C4 English Tokenized Samples
This dataset contains tokenized English samples from the C4 (Colossal Clean Crawled Corpus) dataset for natural language processing (NLP) tasks.
The first 125 000 entries from the en split of allenai/c4
were tokenized using spaCy's en_core_web_sm model. Tokens joined with spaces.
Features
text: Original text from C4
tokenized: The tokenized and space-joined text
num_tokens: Number of tokens after tokenization
num_punct_tokens: Number of… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/c4-en-tokenized.redpajama-c4-refined-by-data-juicer
RedPajama -- C4 (refined by Data-Juicer)
A refined version of C4 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 832GB).
Dataset Information
Number of samples: 344,491,171 (Keep ~94.42% from the original dataset)
Refining Recipe
#… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-c4-refined-by-data-juicer.
