datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.sera-phase2-saudi-dialect-rag
SERA Saudi Dialect RAG Dataset (Phase 2 - Domain)
Domain-specific RAG fine-tuning dataset for Saudi Arabic dialect,
focused on Saudi Electricity Regulatory Authority (SERA) documents.
Format
LlamaFactory Alpaca format:
Field
Description
instruction
System prompt + real document chunk as context + question in Saudi dialect
input
Always empty
output
Answer in Saudi dialect
Usage with LlamaFactory
Copy the JSON files into your LlamaFactory… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/sera-phase2-saudi-dialect-rag.saudi-dialect-rag
Saudi Dialect RAG Fine-Tuning Dataset
A RAG-formatted fine-tuning dataset for Saudi Arabic dialect, built from
HeshamHaroon/saudi-dialect-conversations.
Format
Each example follows the LlamaFactory Alpaca format:
Field
Description
instruction
System prompt + MSA context paragraph + optional conversation history + question
input
Always empty string
output
Assistant reply in Saudi dialect
How it was built
Loaded source multi-turn Saudi… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/saudi-dialect-rag.Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT
Turkish Dialectical Reasoning Dataset (Sokrates-ToT)
The Turkish Dialectical Reasoning Dataset (Sokrates-ToT) is a collection structured in a Tree-of-Thought (ToT) format, based on a multi-persona and dialectical reasoning framework.Inspired by Socrates' method of dialogue, it facilitates deep analysis of complex and multidimensional issues by having AI personas with different expertise interact and ultimately reach a final synthesis.
Purpose of the Dataset
This… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT.Open-ended_Questions_dialectal_data
Dataset Summary
A collection of open-ended questions that was provided to the data marathon competitors to populate KIND dataset. It was designed to elicit longer responses cultural and context-rich sentences.
For more details, please check the paper
The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection
Citation Information
@inproceedings{yamani-etal-2024-kind,
title = "The {KIND} Dataset: A Social Collaboration Approach for Nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KIND-Dataset/Open-ended_Questions_dialectal_data.exam_zh_multitopic_dialect_culture
exam_zh_multitopic_dialect_culture
This dataset contains 300 multiple-choice questions (MCQs) from a variety of Mandarin-based assessments, spanning both regional dialect comprehension and cultural/general knowledge.
📚 Description
The questions come from publicly available Chinese-language exams and quizzes, and fall into two major categories:
🗣️ Regional Dialect Tests
These assess language understanding across major Chinese dialects and topolects:
Hakka… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/exam_zh_multitopic_dialect_culture.hadrami-arabic-dialect-dataset
Hadrami Arabic Dialect Dataset
A structured dataset of 1,000 entries from the Hadrami Arabic dialect (spoken primarily in the Hadramawt region of Yemen). Each entry includes the dialectal word alongside its Modern Standard Arabic (MSA/Fusha) equivalent, linguistic metadata, usage examples, proverbs, and semantic tags.
🔊 Roadmap: Audio pronunciations for each entry are planned for a future release.
Dataset Summary
Field
Value
Entries
1,000
Language… See the full description on the dataset page: https://huggingface.co/datasets/saeedbark/hadrami-arabic-dialect-dataset.east_java_dialect_instruct
Complaints From The East Javanese Dialect community
This dataset created manually by humans with reference to public complaints in the comments column of the local government's Instagram account and another platform like X and TikTok Comments.
Jordanian-Dialect-Instruct-QA
Dataset Card: Jordanian-Dialect-Instruct-QA
Overview
This dataset contains 1500 academic Q&A pairs for Jordanian universities, written in the Jordanian Arabic dialect. It is used to fine-tune models to provide natural, localized responses to academic inquiries and more.
This dataset was used to fine-tune Qwen2.5-7B-Instruct-Jordanian. You can visit the model page to download the GGUF weights and the custom Ollama Modelfile.
Dataset Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/KareemBb/Jordanian-Dialect-Instruct-QA.Understanding_Dialect_Text
🇰🇿 Kazakh Dialect Analysis and Standardization Dataset
Dataset Summary
Kazakh Dialect Analysis and Standardization Dataset is a Kazakh-language linguistics dataset designed for dialect identification, dialectal feature analysis, and normalization into standard Kazakh.
The dataset contains instruction-style prompts asking the model to identify dialectal or regional language features in a given Kazakh text and explain how they can be converted into standard… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Understanding_Dialect_Text.
