nandhakumarms/qualc-fineweb-edu-en
QualC FineWeb-Edu English (Cleaned) QualC FineWeb-Edu English (Cleaned) is a cleaned subset of the official FineWeb-Edu dataset published by Hugging Face. The dataset is intended for Large Language Model (LLM) pretraining, tokenizer training, continual pretraining, educational NLP research, and language modeling. This repository contains approximately one million cleaned educational English documents prepared for the QualC project. Dataset Information Item… See the full description on the dataset page: https://huggingface.co/datasets/nandhakumarms/qualc-fineweb-edu-en.
0131
Update README.md
Update README.md
Upload dataset
initial commit
