CoolFace
Datasetpublic

orionai/en_wikipedia_001

Dataset Card for en_wikipedia_001 The en_wikipedia_001 dataset is a collection of crawled paragraph text from Wikipedia on the 28th of April, 2024. It contains high-quality text, stored in multiple documents, available to be used to finetune or train AI models based that the license is followed. Dataset Details The dataset was crawled using our web crawler on the 28th of April at an average of 1 page per second as to respect robots.txt rules. Strict licensing must… See the full description on the dataset page: https://huggingface.co/datasets/orionai/en_wikipedia_001.

sourceHugging Faceotherupdated 2y agoView on Hugging Face
2likes38downloads

orionai/en_wikipedia_001 · main · files are served by the source, never re-hosted here