CoolFace
Datasetpublic

levantdata/jordanian-dialect-sample-v1

Jordanian Dialect Sample (v1) Levant AI is building dialect-authentic Arabic training data for the Levantine region, starting with Jordanian/Shami dialect — collected natively, not translated from Modern Standard Arabic or English. This is our first public sample: 30 examples spanning three categories that reflect real gaps in current Arabic AI training data: Categories general_conversation (10 examples) — everyday natural Jordanian dialect exchanges… See the full description on the dataset page: https://huggingface.co/datasets/levantdata/jordanian-dialect-sample-v1.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes18downloads
Dataset Card

Jordanian Dialect Sample (v1)

Levant AI is building dialect-authentic Arabic training data for the Levantine region, starting with Jordanian/Shami dialect — collected natively, not translated from Modern Standard Arabic or English.

This is our first public sample: 30 examples spanning three categories that reflect real gaps in current Arabic AI training data:

Categories

  • `general_conversation` (10 examples) — everyday natural Jordanian dialect exchanges
  • `function_calling` (10 examples) — dialect phrasing with clear actionable intent (booking, cancellation, price inquiry, order tracking), addressing the documented accuracy gap in Arabic function-calling
  • `safety_redteaming` (10 examples) — realistic dialect-based prompt injection and social engineering attempts, addressing the documented gap where translated safety guardrails fail on native dialect input

Format

JSONL, one object per line with fields: category, dialect, textarabic, textenglish_gloss, notes

Why this matters

Most Arabic NLP datasets are either Modern Standard Arabic or machine-translated from English, which fails to capture how people actually speak — and fails specifically in safety and tool-use scenarios where dialect nuance changes model behavior.

This dataset is manually written and reviewed by a native Jordanian speaker, not machine-translated.

Roadmap

This is v1 (sample). Planned expansion: 200-500+ examples per category, additional Levantine sub-dialects, and a living/maintained update cycle.

Contact

Built by LevantData. Open to collaboration, feedback, and partnership inquiries.