CoolFace
Datasetpublic

Raziel1234/WebText-3

WebText-3 Corpus WebText-3 is a large-scale, diverse text corpus collected from publicly available web pages. It contains cleaned and normalized sentences suitable for natural language processing (NLP), machine learning, and AI training. Dataset Overview Format: Plain text (.txt), one sentence per line Approximate Size: 200,000+ sentences Languages: Primarily English, with occasional Hebrew content Source Types: Wikipedia articles, technology news sites, blogs… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-3.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes23downloads
7 commits on main
a79e4181y ago

Upload corpus.txt

Raziel1234
e822bc11y ago

Delete corpus.txt

Raziel1234
bf34e3b1y ago

Upload corpus.txt

Raziel1234
9979f4c1y ago

Create LICENSE

Raziel1234
dff651c1y ago

Update .gitattributes

Raziel1234
b5365421y ago

Update README.md

Raziel1234
e6465cf1y ago

initial commit

Raziel1234