orionai/en_wikipedia_001
Dataset Card for en_wikipedia_001 The en_wikipedia_001 dataset is a collection of crawled paragraph text from Wikipedia on the 28th of April, 2024. It contains high-quality text, stored in multiple documents, available to be used to finetune or train AI models based that the license is followed. Dataset Details The dataset was crawled using our web crawler on the 28th of April at an average of 1 page per second as to respect robots.txt rules. Strict licensing must… See the full description on the dataset page: https://huggingface.co/datasets/orionai/en_wikipedia_001.
Update README.md
Upload 0000.csv
Delete 0000.txt
Rename scraped.txt to 0000.txt
Update README.md
Upload scraped.txt
Delete 0000.parquet
Update README.md
Delete data-001.txt
Upload 0000.parquet
Update README.md
Update README.md
Rename scraped.txt to data-001.txt
Upload scraped.txt
Update LICENSE
initial commit
