datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sada2022
Dataset Card for SADA صدى
Dataset Summary
يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر.
ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر مجموعة… See the full description on the dataset page: https://huggingface.co/datasets/khaledalganem/sada2022.ViSL-News
ViSL-News
Dataset Summary
ViSL-News is a sentence-level Vietnamese Sign Language (VSL) dataset constructed from sign-interpreted Vietnamese news broadcasts.
The dataset was built from HTV Tin Tức videos published on YouTube during 2024–2025. Each sample consists of a sentence-level sign-language video clip paired with a Vietnamese text sentence.
ViSL-News was constructed using ViSL-Tool, a semi-automated framework designed for news videos that contain spoken… See the full description on the dataset page: https://huggingface.co/datasets/kha2612/ViSL-News.autonlp-data-CoronaIt's all about Corona
SADA_khaledalganemsada2022_Rawdate
Dataset Card for SADA صدى
Dataset Summary
يعتبر توفر البيانات من أهم ممكنات تطوير نماذج ذكاء اصطناعي متفوقة إن لم يكن أهمها، ولكن لا تزال البيانات الصوتية المفتوحة وخصوصاً باللغة العربية ولهجاتها المختلفة شحيحة المصدر.
ومن هذا المنطلق وحرصًا على إطلاق القيمة الكامنة للبيانات وتمكين تطوير منتجات مبنية على الذكاء الاصطناعي، قام المركز الوطني للذكاء الاصطناعي في سدايا (الهيئة الوطنية للبيانات والذكاء الاصطناعي) بالتعاون مع الهيئة السعودية للإذاعة والتلفزيون بنشر… See the full description on the dataset page: https://huggingface.co/datasets/Sundus246/SADA_khaledalganemsada2022_Rawdate.darja-blindspot-eval
Blind Spot Evaluation: Algerian Darja, Arabizi, and French Code-Switching
Author: Khadija Abderrahmane
Model evaluated: Qwen/Qwen2.5-1.5B-Instruct (1.5B parameters)
1. The blind spot
Nearly every widely used Arabic NLP benchmark — ArabicMMLU, ARLUE, AraSentiment, and most Arabic instruction-tuning datasets — is built almost entirely on Modern Standard Arabic (MSA), with limited coverage of major spoken dialects (Egyptian, Gulf, Levantine). Algerian Darja, the… See the full description on the dataset page: https://huggingface.co/datasets/khadidjaabderrahmane/darja-blindspot-eval.12-con-giap-hop-khac
Quan hệ hợp khắc 12 con giáp
Compatibility matrix of the 12 earthly branches
1. Mô tả · Description
Quan hệ giữa mọi cặp trong 12 địa chi, đủ 144 dòng, xuất dạng dài nên mỗi cặp một dòng.
The relation between every pair of the 12 earthly branches, 144 rows, in long form with one pair per row.
Số dòng · Rows: 144
Phiên bản · Version: 1.0.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure
Cột · Column
Kiểu · Type
Ý nghĩa… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/12-con-giap-hop-khac.AMPS_khankhayyam-challengemuslim-names-dataset
Muslim Names Dataset
A comprehensive collection of Muslim names with meanings scraped from muslimnames.com. Contains 14,585 names with English names, Arabic names, meanings, and gender classifications.
Dataset Contents
This dataset contains ~14,585 Muslim names with the following information:
English name: Name in English/Latin script
Arabic name: Name in Arabic script
Meaning: Definition and meaning of the name
Gender: Classification as male or female… See the full description on the dataset page: https://huggingface.co/datasets/khalidAboubakr/muslim-names-dataset.forcis
Dataset Card for Processed FORCIS Data
Dataset Name: Processed FORCIS Data
Source: FRBCesab/forcis (originally from Zenodo: https://zenodo.org/records/12724286)
Description:
This dataset is a processed version of the FORCIS data, obtained from Zenodo. The original data uses a custom delimiter (;) and may contain leading/trailing whitespace and special characters. This processed version addresses these issues and provides a clean, comma-separated values (CSV) format for easier use.… See the full description on the dataset page: https://huggingface.co/datasets/khammami/forcis.MBIB
Dataset Description
This dataset contains news articles annotated with political bias labels such as left, center, and right. The labels are derived from the political orientation of the news source, based on assessments by the Media Bias/Fact Check (MBFC) project. It is designed for training and evaluating transformer-based models on the task of ideological bias classification in news content.
bhasaflow-khasi-monolingual-corpus-v1
BhasaFlow Khasi Monolingual Corpus v1
By Medharvix Systems Private Limited
Overview
A curated monolingual Khasi text corpus for language modeling, NLP research, and linguistic analysis, with a focus on preserving and digitizing low-resource languages of Northeast India.
Dataset Structure
Column
Description
khasi_sentence
Khasi language sentence
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-monolingual-corpus-v1.bhasaflow-khasi-english-parallel-corpus-v1
BhasaFlow Khasi-English Parallel Corpus v1
By Medharvix Systems Private Limited
Overview
A curated parallel corpus of Khasi-English sentence pairs designed for machine translation research and development, with a focus on low-resource language technology for Northeast India.
Dataset Structure
Column
Description
sentence_id
Unique sentence identifier
english_text
English sentence
khasi_text
Khasi translation
Usage
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-corpus-v1.FEM-Khasi-News-Monolingual-Corpus
Khasi Monolingual News Corpus (740K)
Project Attribution & Collaboration
This dataset was collected and curated as part of the research project titled "Financial Empowerment in Meghalaya: AI-Powered Multilingual E-Marketplace for Tribes."
This project is a collaborative research initiative conducted by:
National Law University (NLU) Meghalaya
Indian Institute of Information Technology (IIIT) Guwahati
Contributors:
This dataset is the result of a joint effort by the… See the full description on the dataset page: https://huggingface.co/datasets/Bapynshngain/FEM-Khasi-News-Monolingual-Corpus.khasi-more-rawEnglish-Khasi-Parallel-Corpus-v1data_source:
Web-scraped data manually vetted by me.
NIT Silchar’s "EnKhCorp1.0: An English–Khasi Corpus."
Tatoeba project.
samanantar : the largest publicly available parallel corpora collection for 11 indic languages
acknowledgments:
Special thanks to Ahlad, NIT Silchar for their "EnKhCorp1.0" dataset, IIT Madras and other institutions involved in the creation of Samanantar and to the contributors of the Tatoeba project.
german-cities-open-data
InfraNode German Cities Open-Data Snapshot
Ein reproduzierbarer, offen lizenzierter Querschnitt von Infrastruktur- und
Umweltdaten für 84+ deutsche Städte, erzeugt aus der öffentlichen
InfraNode-API. Eine Zeile je Stadt.
Inhalt
Bereich
Felder
Quelle
Stammdaten
slug, name_de, state, ags, wikidata_qid, lat, lon, base_population, base_area_km2
Wikidata (CC0)
Wetter
weather_temperature_c, weather_humidity, weather_condition
DWD (GeoNutzV)
Luftqualität… See the full description on the dataset page: https://huggingface.co/datasets/Khaledc83/german-cities-open-data.stellar_classificationarabigee-data
ArabiGEE Data
Paper title: ArabiGEE: A Hierarchical Taxonomy for Arabic Grammatical Error Explanation
annotations.csv
context_id: Links the annotation row to either context file.
pair_id: Word-pair identifier in this dataset.
original_pair_id: Original pair identifier in the source data.
error_id: Error number within the same pair_id.
erroneous_word: Erroneous word or phrase.
target_word: Target word or phrase.
areta_label: ARETA edit label.
lex_code: Lexical… See the full description on the dataset page: https://huggingface.co/datasets/khaled44/arabigee-data.code-switching-codesaviours-si26-zainab
Roman Urdu–English Code-Switching Dataset
Dataset Description
This dataset contains 1,901 sentences and 21,370 word-level language labels, built to capture how Roman Urdu and English are naturally mixed together in everyday Pakistani online communication.
Code-switching — blending two languages within a single sentence — is how the vast majority of Pakistanis actually write and speak online, on platforms like Twitter/X, Facebook, YouTube, Reddit, and WhatsApp. A… See the full description on the dataset page: https://huggingface.co/datasets/Zainab-Binte-Khalid/code-switching-codesaviours-si26-zainab.khasi-essayscauhoilichsukhasi-datasets
What is Khasi Language?
Location:
Primarily spoken in the northeastern Indian state of Meghalaya.
Also spoken in parts of Assam, Tripura, and Bangladesh.
Language Family:
Khasi is a member of the Austroasiatic language family.
Script:
Traditionally written using the Khasi script, which is a script created specifically for the Khasi language.
Culture and Identity:
The Khasi language is an integral part of the cultural identity of the… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/khasi-datasets.Fixberry
FixBerry
Fixberry is a little dataset I have made to train models to correctly count the number of letters in a word. It is commonly known that even the best LLMs fail at counting the number of R's in strawberry. I have also found out they have problems with other words too, like keeper and parallel but weirdly not with words like pepper and peeper. This should really be investigated more closely, I suspect it has something to do with tokenization and the possability that the model… See the full description on the dataset page: https://huggingface.co/datasets/Khawn2u/Fixberry.osworld_tasks_filesKhavee-klonkhamenei_ir_1352_1403_08_13QG_pythonmovie-quotes1movie-quotes
