amanpreet7/allenai-c4
🧠 ALLENAI C4 - English Train Split (Prepared Version) This repository contains the preprocessed and ready-to-use version of the ALLENAI C4 (Colossal Clean Crawled Corpus) English train split. It has been downloaded and optionally transformed for downstream NLP tasks such as pretraining large language models or text-based retrieval systems. 📦 Dataset Details Original Source: allenai/c4 Language: English (en) Split: train License: Google C4 License ⚠️ Note: This version only includes the train… See the full description on the dataset page: https://huggingface.co/datasets/amanpreet7/allenai-c4.
🧠 ALLENAI C4 - English Train Split (Prepared Version) This repository contains the preprocessed and ready-to-use version of the ALLENAI C4 (Colossal Clean Crawled Corpus) English train split. It has been downloaded and optionally transformed for downstream NLP tasks such as pretraining large language models or text-based retrieval systems.
📦 Dataset Details Original Source: allenai/c4
Language: English (en)
Split: train
License: Google C4 License
⚠️ Note: This version only includes the train split and is fully preprocessed & stored locally (e.g., in Apache Arrow or Parquet format)
📊 Size Info Number of samples: ~365 million
Storage format: Arrow
Total disk size: ~881GB(approx)
