Raziel1234/WebText-3
WebText-3 Corpus WebText-3 is a large-scale, diverse text corpus collected from publicly available web pages. It contains cleaned and normalized sentences suitable for natural language processing (NLP), machine learning, and AI training. Dataset Overview Format: Plain text (.txt), one sentence per line Approximate Size: 200,000+ sentences Languages: Primarily English, with occasional Hebrew content Source Types: Wikipedia articles, technology news sites, blogs… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-3.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face