datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sada2022
Dataset Card for SADA صدى
Dataset Summary
يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر.
ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر مجموعة… See the full description on the dataset page: https://huggingface.co/datasets/khaledalganem/sada2022.ViSL-News
ViSL-News
Dataset Summary
ViSL-News is a sentence-level Vietnamese Sign Language (VSL) dataset constructed from sign-interpreted Vietnamese news broadcasts.
The dataset was built from HTV Tin Tức videos published on YouTube during 2024–2025. Each sample consists of a sentence-level sign-language video clip paired with a Vietnamese text sentence.
ViSL-News was constructed using ViSL-Tool, a semi-automated framework designed for news videos that contain spoken… See the full description on the dataset page: https://huggingface.co/datasets/kha2612/ViSL-News.autonlp-data-CoronaIt's all about Corona
SADA_khaledalganemsada2022_Rawdate
Dataset Card for SADA صدى
Dataset Summary
يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر.
ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر… See the full description on the dataset page: https://huggingface.co/datasets/Sundus246/SADA_khaledalganemsada2022_Rawdate.darja-blindspot-eval
Blind Spot Evaluation: Algerian Darja, Arabizi, and French Code-Switching
Author: Khadija Abderrahmane
Model evaluated: Qwen/Qwen2.5-1.5B-Instruct (1.5B parameters)
1. The blind spot
Nearly every widely used Arabic NLP benchmark — ArabicMMLU, ARLUE, AraSentiment, and most Arabic instruction-tuning datasets — is built almost entirely on Modern Standard Arabic (MSA), with limited coverage of major spoken dialects (Egyptian, Gulf, Levantine). Algerian Darja, the… See the full description on the dataset page: https://huggingface.co/datasets/khadidjaabderrahmane/darja-blindspot-eval.khayyam-challengeforcis
Dataset Card for Processed FORCIS Data
Dataset Name: Processed FORCIS Data
Source: FRBCesab/forcis (originally from Zenodo: https://zenodo.org/records/12724286)
Description:
This dataset is a processed version of the FORCIS data, obtained from Zenodo. The original data uses a custom delimiter (;) and may contain leading/trailing whitespace and special characters. This processed version addresses these issues and provides a clean, comma-separated values (CSV) format for easier use.… See the full description on the dataset page: https://huggingface.co/datasets/khammami/forcis.MBIB
Dataset Description
This dataset contains news articles annotated with political bias labels such as left, center, and right. The labels are derived from the political orientation of the news source, based on assessments by the Media Bias/Fact Check (MBFC) project. It is designed for training and evaluating transformer-based models on the task of ideological bias classification in news content.
german-cities-open-data
InfraNode German Cities Open-Data Snapshot
Ein reproduzierbarer, offen lizenzierter Querschnitt von Infrastruktur- und
Umweltdaten für 84+ deutsche Städte, erzeugt aus der öffentlichen
InfraNode-API. Eine Zeile je Stadt.
Inhalt
Bereich
Felder
Quelle
Stammdaten
slug, name_de, state, ags, wikidata_qid, lat, lon, base_population, base_area_km2
Wikidata (CC0)
Wetter
weather_temperature_c, weather_humidity, weather_condition
DWD (GeoNutzV)
Luftqualität… See the full description on the dataset page: https://huggingface.co/datasets/Khaledc83/german-cities-open-data.stellar_classificationarabigee-data
ArabiGEE Data
Paper title: ArabiGEE: A Hierarchical Taxonomy for Arabic Grammatical Error Explanation
annotations.csv
context_id: Links the annotation row to either context file.
pair_id: Word-pair identifier in this dataset.
original_pair_id: Original pair identifier in the source data.
error_id: Error number within the same pair_id.
erroneous_word: Erroneous word or phrase.
target_word: Target word or phrase.
areta_label: ARETA edit label.
lex_code: Lexical… See the full description on the dataset page: https://huggingface.co/datasets/khaled44/arabigee-data.osworld_tasks_filesgpt89-v2mmv1qwen-reasondistill-logs50_startupCovid
COVID-19 Patient Symptoms Dataset
Overview
This dataset contains patient-level records related to COVID-19 symptoms and diagnosis outcomes. Each record captures a combination of demographic information, clinical indicators, and a final diagnosis label.
The data spans multiple Indian cities, providing a diverse sample across different age groups and symptom intensities.
Features Included
Demographic Details
Age
Gender
City
Clinical Indicators
Fever… See the full description on the dataset page: https://huggingface.co/datasets/KHALIFAHARBY/Covid.toxic_comentstellar_classification_missingRecipe_Generation_Dataset
