datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sada2022
Dataset Card for SADA صدى
Dataset Summary
يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر.
ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر مجموعة… See the full description on the dataset page: https://huggingface.co/datasets/khaledalganem/sada2022.sada_clean_envSADA_khaledalganemsada2022_Rawdate
Dataset Card for SADA صدى
Dataset Summary
يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر.
ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر… See the full description on the dataset page: https://huggingface.co/datasets/Sundus246/SADA_khaledalganemsada2022_Rawdate.sada2022-arabic-tts
SADA 2022 - Saudi Arabic Dataset for TTS
مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech)
المصدر الأصلي
Kaggle: sdaiancai/sada2022
الاستخدام
# طريقة 1: Git Clone
!git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset
# طريقة 2: مكتبة datasets
from datasets import load_dataset
dataset = load_dataset("aalshalfi/sada2022-arabic-tts")
الملفات
valid.csv - ملف البيانات الرئيسي
wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts.sa-data
Storia dell'Arte Dataset (SA-Data)
📌 Descrizione del Dataset
Il dataset SA-Data è una raccolta strutturata di articoli della rivista Storia dell'Arte (https://www.storiadellarterivista.it/) digitalizzati e arricchiti con metadati dettagliati e rappresentazioni semantiche. È stato creato per supportare la ricerca accademica e le applicazioni di elaborazione del linguaggio naturale.
🔍 Contenuto
Il dataset include:
1050 articoli pubblicati tra il… See the full description on the dataset page: https://huggingface.co/datasets/paolodegasperis/sa-data.sadai-mrec-query-rewrite-13karabic_eou_sada_dataset
Arabic EOU SADA Dataset (Saudi Dialect)
414,053 conversational Arabic utterances annotated for End-of-Utterance (EOU) detectionStrong focus on natural Saudi dialect (خليجي / نجدي / حجازي)
Task
Binary classification:
label = 1 → End of speaker turn (EOU)
label = 0 → Speaker will continue
Columns
text: Arabic transcription
label: 0 or 1
silence_after_seconds: pause duration after this segment
split: train | validation | test (already included)… See the full description on the dataset page: https://huggingface.co/datasets/LordTenson/arabic_eou_sada_dataset.11taxon-est-gtavian-msa-iqtreesada-eou-saudi-dialectsa-data
Storia dell'Arte Dataset (SA-Data)
📌 Descrizione del Dataset
Il dataset SA-Data è una raccolta strutturata di articoli della rivista Storia dell'Arte (https://www.storiadellarterivista.it/) digitalizzati e arricchiti con metadati dettagliati e rappresentazioni semantiche. È stato creato per supportare la ricerca accademica e le applicazioni di elaborazione del linguaggio naturale.
🔍 Contenuto
Il dataset include:
1050 articoli pubblicati tra il 1969 e… See the full description on the dataset page: https://huggingface.co/datasets/jacobdc1234/sa-data.sada-diarization-preview
SADA 2022 Arabic Diarization
Training-ready speaker-attributed ASR windows derived from
SADA 2022. The source
recordings are mirrored at
khaledalganem/sada2022.
Splits
train: 36,004 windows, 202.064 hours, 4,062 recordings
validation: 853 windows, 4.774 hours, 88 recordings
test: 901 windows, 5.006 hours, 111 recordings
Total: 37,758 windows and
211.844 hours.
The official SADA train, validation, and test partitions are preserved.
Windows are 8–28 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Mohaddz/sada-diarization-preview.clinical-nlp-patient-notesSpeed.Datasetsada2022-eousada2022-arabic-tts
SADA 2022 - Saudi Arabic Dataset for TTS
مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech)
المصدر الأصلي
Kaggle: sdaiancai/sada2022
الاستخدام
# طريقة 1: Git Clone
!git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset
# طريقة 2: مكتبة datasets
from datasets import load_dataset
dataset = load_dataset("aalshalfi/sada2022-arabic-tts")
الملفات
valid.csv - ملف البيانات الرئيسي
wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/mohamedalgmaal/sada2022-arabic-tts.clinical-nlp-featuresarabic_eou_sada_curated11taxon-true-gtmammalian-424genesclinical-nlp-trainclinical-nlp-datasetclinical-nlp-dataset-raw
Clinical NLP Dataset (Raw)
Includes: train.csv, test.csv, patient_notes.csv, features.csv
clinical-nlp-testSADA_EOU
