datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
saudi1saudi-english-code-switching-datasetAqar.fm-Saudi-Real-Estate-Listings
Aqar.fm Saudi Real Estate Listings Dataset
Dataset Description
This dataset contains real estate listings scraped from sa.aqar.fm, covering various regions in Saudi Arabia. It includes data for sales, rentals, and auctions.
Total Records: 6,252
Date Collected: 30/12/2025
Source: Public listings on sa.aqar.fm
File Structure
The dataset is provided in CSV, JSON, and Notebook formats:
aqar_fm_listings_cleaned.csv: The main dataset containing all cleaned… See the full description on the dataset page: https://huggingface.co/datasets/afaskar/Aqar.fm-Saudi-Real-Estate-Listings.SaudiTraditionalFoodAugmentedThis dataset is for our paper entitled "Saudi Traditional Food Recognition Using Deep Learning."
Abstract:
The use of deep learning for food recognition has attracted significant attention due to its applications in dietary monitoring, automated food logging, and nutrition analysis. While deep learning has demonstrated impressive results across various cuisines, some, including Saudi Arabian cuisine, have not been thoroughly explored. This paper examines the effectiveness of deep learning… See the full description on the dataset page: https://huggingface.co/datasets/ssalahmari/SaudiTraditionalFoodAugmented.ipfs_saudiarabia_laws_ir
Saudi Arabia legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_saudiarabia_laws (revision 94bd316e19d5d391c4b9e202d199d5debf8965c3) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Saudi Arabia prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_saudiarabia_laws_ir.Lahgtna-saudi
Lahgtna Saudi (Mans1611/Lahgtna-saudi)
Saudi-dialect subset prepared for Arabic ASR fine-tuning
(from oddadmix/dialectal-arabic-lahgtna-v2, filtered to language == "sa").
Splits
Split
Rows
train
11,030
test
581
Columns
audio, text, language, duration
Features
{'audio': Audio(sampling_rate=16000, decode=False, num_channels=None, stream_index=None), 'text': Value('string'), 'language': Value('string'), 'duration':… See the full description on the dataset page: https://huggingface.co/datasets/Mans1611/Lahgtna-saudi.saudi2saudi-dialect-conversations
Saudi Najdi Dialect Conversations
A curated dataset of 3,545 multi-turn conversations in Saudi Najdi Arabic dialect (the dialect spoken in Riyadh, Qassim, and central Najd region). Designed for Supervised Fine-Tuning (SFT) of Arabic language models.
Dataset Details
Metric
Value
Total conversations
3,545
Total turns
22,536
Average turns per conversation
6.4
Complexity distribution
Simple: 31%, Intermediate: 38%, Advanced: 31%
Topics covered
18 categories… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/saudi-dialect-conversations.ALL_WSP_Saudiarabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.saudi_dialect_asrv1.0ipfs_saudiarabia_laws
Saudi Arabia In-Force Laws (Bureau of Experts / BOE)
Research snapshot of official national legislation from Bureau of Experts at the Council of Ministers (Wayback of official laws.boe.gov.sa URLs).
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-18
Coverage
wayback-of-official snapshot
Source
Bureau of Experts at the Council of Ministers (Wayback of official… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_saudiarabia_laws.saudi-dialect-speech-female
🌍 Saudi Dialectal Arabic Audio Dataset
This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning.
🗂️ Dataset Columns
Column
Description
audio
The audio chunk (22,050 Hz, mono WAV)
duration
Chunk duration in seconds
base_transcription
Transcript from the base Arabic ASR model
dialectal_transcription
Transcript from the Saudi-dialectal… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/saudi-dialect-speech-female.saudi2_conSaudi-Arabia-Stock-Symbols-and-Metadata
Saudi Arabia Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Saudi Arabia.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Saudi-Arabia-Stock-Symbols-and-Metadata.saudinewsnetThe dataset contains a set of 31,030 Arabic newspaper articles alongwith metadata, extracted from various online Saudi newspapers and written in MSA.saudi1_con_tempsaudi-tts-synthetic-200ksaudi-arabic-laws-and-regulations-corpus
Saudi Arabic Laws and Regulations Corpus
A structured, article-level Arabic legal corpus containing 22,593 records from 578 official Saudi legal documents, prepared for Arabic legal information retrieval, Retrieval-Augmented Generation (RAG), grounded generation, and LLM evaluation.
Release: 1.0.0Language: ArabicDomain: Saudi laws and regulationsGranularity: Legal articleTotal records: 22,593Archival DOI: 10.5281/zenodo.21265180
Quick Start
The corpus is… See the full description on the dataset page: https://huggingface.co/datasets/SaudiArabicLaws/saudi-arabic-laws-and-regulations-corpus.saudi_up
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/yeeaee/saudi_up.SaudiTalk
SaudiTalk: A Multi-Source Dialectal Speech Dataset from Saudi Arabia
Dataset Summary
SaudiTalk is a curated and human-verified Arabic speech dataset covering three major Saudi dialects: Hijazi, Ha’il, and Southern. The dataset is constructed from publicly available social media content and is designed to support research in Automatic Speech Recognition (ASR), dialect identification, and Arabic speech processing.
Key Features
3 Saudi dialects: Hijazi, Ha’il… See the full description on the dataset page: https://huggingface.co/datasets/SaudiTalk/SaudiTalk.30k-SADA22_Saudisaudipedia-arabic-qa
Saudipedia Q&A Dataset
Dataset Description
Summary
This dataset contains question-answer pairs scraped from Saudipedia, a comprehensive Arabic encyclopedia focused on Saudi Arabia. The dataset includes 1,082 Q&A entries covering various topics related to Saudi culture, history, economy, government, society, geography, religion, and notable personalities.
The data was collected by scraping the website's question-answer section, which provides detailed answers to… See the full description on the dataset page: https://huggingface.co/datasets/AhmadHakami/saudipedia-arabic-qa.saudi-dialect-speech-male268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample
Description
Arabic(Saudi) Multi-stream Spontaneous Dialogue Smartphone speech dataset-Customer Service. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(268 native speakers), geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1627?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample.Saudi_Traditional_FoodSaudi Traditional Food Recognition Using Deep
Learning
Abstract:
The use of deep learning for food recognition has attracted signifi-
cant attention due to its applications in dietary monitoring, automated
food logging, and nutrition analysis. While deep learning has demon-
strated impressive results across various cuisines, some, including Saudi
Arabian cuisine, have not been thoroughly explored. This paper exam-
ines the effectiveness of deep learning models in accurately classifying and… See the full description on the dataset page: https://huggingface.co/datasets/ssalahmari/Saudi_Traditional_Food.60H-SADA22-Saudisaudi-arabic-cs-conversations
Saudi Arabic Customer Service Conversations — Free 100 Sample
100 synthetic multi-turn conversations in authentic Saudi Arabic dialects
Built for LLM fine-tuning, chatbot training, and Arabic NLP research
Overview
This is a free 100-conversation sample from a production-quality dataset of 50,000 Saudi Arabic customer service conversations. Every conversation is fully synthetic — no real user data — and safe for commercial use.
Each conversation simulates a… See the full description on the dataset page: https://huggingface.co/datasets/dev-hussein/saudi-arabic-cs-conversations.os-rfodg-outdoor-uav-synthetic-dataset-taif-saudi-arabia
UAV Trajectory Simulation Dataset for Terrain-Based Localization
Dataset Overview
This dataset contains simulated UAV flight data generated using ROS2, Gazebo, and PX4 autopilot system. The dataset features a quadcopter performing autonomous flight trajectories over realistic terrain imported from satellite imagery and Digital Elevation Model (DEM) maps of the Taif region in Saudi Arabia.
Dataset Files
The dataset contains:
7 trajectory CSV files:… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/os-rfodg-outdoor-uav-synthetic-dataset-taif-saudi-arabia.saudi-arabia-business-dataset
Saudi Arabia Business Email List — Market Intelligence Dataset
100,000 verified Saudi Arabia contacts are available from LeadsBlue →. This open dataset provides the aggregate market intelligence behind that database — contact volume, benchmark open/reply rates, send timing, and compliance for the Saudi Arabia segment.
At a glance: A research dataset describing the Saudi Arabia Business Email List market: verified-contact volume, industry distribution, outreach benchmarks, and… See the full description on the dataset page: https://huggingface.co/datasets/emailmarketingdataset/saudi-arabia-business-dataset.
