O96a/sudanese-synthetic-instructions
0
Synthetic Sudanese Arabic Instruction Data Generator
Experiment exp-004 | Sudanese NLP domain (PRIORITY) | April 12, 2026
Overview
This experiment tests structured synthetic data generation for Sudanese Arabic instruction following, inspired by the "Agent-as-Annotators" paper's approach to structured trajectory synthesis.
Hypothesis
Structured synthetic dialogue generation using seed patterns can create viable instruction-following data for Sudanese Arabic, even with minimal authentic examples.
Method
- Seed with authentic Sudanese Arabic conversational patterns
- Apply template-based variation
- Generate instruction-response pairs across categories
- Compare with MSA equivalents
Categories Covered
- Greetings: "هاي كيفك؟", "صباح الخير", etc.
- Daily Tasks: "شنو بتعمل النهارده؟", "وين رايح؟"
- Requests: "ممكن تجيب لي مي؟", "عندك فكة؟"
- Food: "شنو الغدا؟", "جعان كتير"
Key Findings
Sudanese vs MSA Differences
Synthetic Data Quality
- Lexical diversity: Limited by seed pattern diversity
- Dialect authenticity: High for common phrases, needs expansion
- Scalability: Template-based approach scales linearly
- Validation needed: Requires native speaker evaluation
Applications
- Bootstrap fine-tuning for Sudanese chatbots
- Create evaluation benchmarks
- Study dialect-specific instruction following
- Augment limited authentic datasets
Files
app.py: Gradio interface for data generationrequirements.txt: Pinned dependenciesREADME.md: This file
Next Steps
- [ ] Collect real Sudanese conversations for validation
- [ ] Fine-tune small LLM (1B-3B) on synthetic data
- [ ] Evaluate against native speaker judgments
- [ ] Expand pattern library with more categories
References
- Agent-as-Annotators paper: https://huggingface.co/papers/2604.07776
- Related: exp-003 (Sudanese Arabic OCR Stress Test)
Space
https://huggingface.co/spaces/O96a/sudanese-synthetic-instructions
Build trigger
Rebuild trigger Sun Apr 12 18:02:07 UTC 2026
Sun Apr 12 20:40:20 UTC 2026
