Raziel1234/WebText-3
WebText-3 Corpus WebText-3 is a large-scale, diverse text corpus collected from publicly available web pages. It contains cleaned and normalized sentences suitable for natural language processing (NLP), machine learning, and AI training. Dataset Overview Format: Plain text (.txt), one sentence per line Approximate Size: 200,000+ sentences Languages: Primarily English, with occasional Hebrew content Source Types: Wikipedia articles, technology news sites, blogs… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-3.
023
Upload corpus.txt
Delete corpus.txt
Upload corpus.txt
Create LICENSE
Update .gitattributes
Update README.md
initial commit
