QCRI/Aslema-Synth-TN
Aslema-Synth-TN Aslema-Synth-TN is a fully synthetic Tunisian Derja corpus for spoken language understanding: speech annotated for intent and for slot filling. It was built for NADI 2026 Shared Task 5 to cover the intents and slots that are rare or absent in the real SLURP-TN training split, and it is the augmentation set behind the Aslema system, which ranked 1st in slot filling on the official test set. No human was recorded for this dataset. An LLM wrote the utterance text, a… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/Aslema-Synth-TN.
2252
Add parquet shards (audio + annotations, 12 shards)
Add flat metadata table
Add pipeline figure
Add release build script
Add dataset card
initial commit
