CoolFace
Datasetpublic

orionai/en_wikipedia_001

Dataset Card for en_wikipedia_001 The en_wikipedia_001 dataset is a collection of crawled paragraph text from Wikipedia on the 28th of April, 2024. It contains high-quality text, stored in multiple documents, available to be used to finetune or train AI models based that the license is followed. Dataset Details The dataset was crawled using our web crawler on the 28th of April at an average of 1 page per second as to respect robots.txt rules. Strict licensing must… See the full description on the dataset page: https://huggingface.co/datasets/orionai/en_wikipedia_001.

sourceHugging Faceotherupdated 2y agoView on Hugging Face
2likes38downloads
16 commits on main
a6218642y ago

Update README.md

Oscar Wang
04541a42y ago

Upload 0000.csv

Oscar Wang
ea640f52y ago

Delete 0000.txt

Oscar Wang
13b115a2y ago

Rename scraped.txt to 0000.txt

Oscar Wang
b57bcc82y ago

Update README.md

Oscar Wang
250bfa52y ago

Upload scraped.txt

Oscar Wang
34e7fab2y ago

Delete 0000.parquet

Oscar Wang
0e5282b2y ago

Update README.md

Oscar Wang
51cf98e2y ago

Delete data-001.txt

Oscar Wang
c3bd27e2y ago

Upload 0000.parquet

Oscar Wang
72202402y ago

Update README.md

Oscar Wang
3f474502y ago

Update README.md

Oscar Wang
e2f15f12y ago

Rename scraped.txt to data-001.txt

Oscar Wang
c572eba2y ago

Upload scraped.txt

Oscar Wang
5ebc2b62y ago

Update LICENSE

Oscar Wang
12f925a2y ago

initial commit

Oscar Wang