CoolFace
Datasetpublic

QCRI/Aslema-Synth-TN

Aslema-Synth-TN Aslema-Synth-TN is a fully synthetic Tunisian Derja corpus for spoken language understanding: speech annotated for intent and for slot filling. It was built for NADI 2026 Shared Task 5 to cover the intents and slots that are rare or absent in the real SLURP-TN training split, and it is the augmentation set behind the Aslema system, which ranked 1st in slot filling on the official test set. No human was recorded for this dataset. An LLM wrote the utterance text, a… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/Aslema-Synth-TN.

sourceHugging Facecc-by-nc-sa-4.0updated 20d agoView on Hugging Face
2likes252downloads
6 commits on main
883b38320d ago

Add parquet shards (audio + annotations, 12 shards)

hunzed
d18b21520d ago

Add flat metadata table

hunzed
f417fec20d ago

Add pipeline figure

hunzed
4b3db9020d ago

Add release build script

hunzed
2c6ffd220d ago

Add dataset card

hunzed
fc296fb20d ago

initial commit

hunzed