CoolFace
Datasetpublic

amanpreet7/allenai-c4

🧠 ALLENAI C4 - English Train Split (Prepared Version) This repository contains the preprocessed and ready-to-use version of the ALLENAI C4 (Colossal Clean Crawled Corpus) English train split. It has been downloaded and optionally transformed for downstream NLP tasks such as pretraining large language models or text-based retrieval systems. 📦 Dataset Details Original Source: allenai/c4 Language: English (en) Split: train License: Google C4 License ⚠️ Note: This version only includes the train… See the full description on the dataset page: https://huggingface.co/datasets/amanpreet7/allenai-c4.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes888downloads
Dataset Card

🧠 ALLENAI C4 - English Train Split (Prepared Version) This repository contains the preprocessed and ready-to-use version of the ALLENAI C4 (Colossal Clean Crawled Corpus) English train split. It has been downloaded and optionally transformed for downstream NLP tasks such as pretraining large language models or text-based retrieval systems.

📦 Dataset Details Original Source: allenai/c4

Language: English (en)

Split: train

License: Google C4 License

⚠️ Note: This version only includes the train split and is fully preprocessed & stored locally (e.g., in Apache Arrow or Parquet format)

📊 Size Info Number of samples: ~365 million

Storage format: Arrow

Total disk size: ~881GB(approx)