CoolFace
Datasetpublic

Raziel1234/WebText-1

Orion-Spark-30M Dataset The Orion-Spark-30M dataset is a curated corpus containing 10,846 lines of text gathered from reputable internet sources, including Wikipedia pages, technology news websites, and educational platforms. The dataset focuses on foundational and advanced topics related to artificial intelligence, machine learning, large language models, and generative pretrained transformers. It also covers major technology companies, influential figures, and key concepts in… See the full description on the dataset page: https://huggingface.co/datasets/Raziel1234/WebText-1.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes38downloads
9 commits on main
49c35a21y ago

Upload corpus.txt

Raziel1234
31cb5a31y ago

Delete corpus.txt

Raziel1234
94613621y ago

Create LICENSE

Raziel1234
695c93e1y ago

Update README.md

Raziel1234
e6a50dc1y ago

Update README.md

Raziel1234
5481cbe1y ago

Update README.md

Raziel1234
1ce311e1y ago

Update README.md

Raziel1234
cb1047c1y ago

Upload corpus.txt

Raziel1234
ccfe8a01y ago

initial commit

Raziel1234