datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SANAD
Dataset Card for SANAD
Dataset Summary
SANAD Dataset is a large collection of Arabic news articles that can be used in different Arabic NLP tasks such as Text Classification and Word Embedding. The articles were collected using Python scripts written specifically for three popular news websites: AlKhaleej, AlArabiya and Akhbarona. All datasets have seven categories [Culture, Finance, Medical, Politics, Religion, Sports and Tech], except AlArabiya which doesn’t have… See the full description on the dataset page: https://huggingface.co/datasets/khalidalt/SANAD.sanad-fullsanadimdbssana_dataset
Data Collection
Persian domains were collected from the public web and manually reviewed by human annotators. Each domain was categorized based on its primary topic.
After the annotation phase, the domains were crawled using a high-speed distributed crawler. The downloaded webpages were processed in parallel by multiple extraction pipelines.
The crawling system stores metadata related to each page, including crawl timestamps, parent links, domain information, and page depth.… See the full description on the dataset page: https://huggingface.co/datasets/lioradCo/sana_dataset.
