datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sanad_experimentsSanad-ar-datasetSANAD
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/arbml/SANAD.Sanadset650k_sanadset
Sanadset 650K: Data on Hadith Narrators
Dataset Description
Sanadset is a large-scale dataset containing over 650,986 Hadith records collected from 926 historical Arabic books. This dataset was created to assist in the computational analysis of Islamic Hadiths, specifically focusing on the chain of narrators (Sanad) and the content (Matn).
It allows researchers to apply Machine Learning and NLP techniques to tasks such as:
Classifying Hadiths (Strong/Weak).
Analyzing… See the full description on the dataset page: https://huggingface.co/datasets/freococo/650k_sanadset.SANAD
Dataset Card for SANAD
Dataset Summary
SANAD Dataset is a large collection of Arabic news articles that can be used in different Arabic NLP tasks such as Text Classification and Word Embedding. The articles were collected using Python scripts written specifically for three popular news websites: AlKhaleej, AlArabiya and Akhbarona. All datasets have seven categories [Culture, Finance, Medical, Politics, Religion, Sports and Tech], except AlArabiya which doesn’t have… See the full description on the dataset page: https://huggingface.co/datasets/khalidalt/SANAD.sanad-fullSANAD
Arabic News Articles Dataset
About Dataset
Context
SANAD Dataset is a large collection of Arabic news articles that can be used in different Arabic NLP tasks such as Text Classification and Word Embedding. The articles were collected using Python scripts written specifically for three popular news websites: AlKhaleej, AlArabiya and Akhbarona.
All datasets have seven categories [Culture, Finance, Medical, Politics, Religion, Sports and Tech], except AlArabiya which doesn’t… See the full description on the dataset page: https://huggingface.co/datasets/Mouwiya/SANAD.sanadoptuna-logs-task-16sana_data_publicHeart-Disease-Prediction-datasetsemrelimdbsarabic-SANAD-5k-sampleoptuna-logs-task-17sanad_experimentsoptuna-logs-task-22sanad_dfoptuna-logs-task-21roots_ar_sanadROOTS Subset: roots_ar_sanad
sanad
Dataset uid: sanad
Description
Homepage
Licensing
Speaker Locations
Sizes
0.1312 % of total
1.2094 % of ar
BigScience processing steps
Filters applied to: ar
dedup_document
dedup_template_soft
filter_remove_empty_docs
remove_html_spans_sanad
filter_small_docs_bytes_300
arabic-SANAD-religion-sampleArabicEmpatheticDialogues-with-QuranFAQs-for-SMEsoptuna-logs-task-19sana_dataset
Data Collection
Persian domains were collected from the public web and manually reviewed by human annotators. Each domain was categorized based on its primary topic.
After the annotation phase, the domains were crawled using a high-speed distributed crawler. The downloaded webpages were processed in parallel by multiple extraction pipelines.
The crawling system stores metadata related to each page, including crawl timestamps, parent links, domain information, and page depth.… See the full description on the dataset page: https://huggingface.co/datasets/lioradCo/sana_dataset.ELQVSMEs-datasetoptuna-logs-task-20medquad
