datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian_Derja_Dataset
Tunisian Derja Dataset (Tunisian Arabic)
This is a collection of Tunisian dialect textual documents for Language Modeling.
It was used for continual pre-training of Jais LLM model to Tunisian dialect
Languages
The dataset is available in Tunisian Arabic (Derja).
Data Table
subset
Lines
Tunisian_Dialectic_English_Derja
1220712
HkayetErwi
966
Derja_tunsi
13037
TunBERT
67219
TunSwitchCodeSwitching
394163
TunSwitchTunisiaOnly
380546
Tweet_TN… See the full description on the dataset page: https://huggingface.co/datasets/linagora/Tunisian_Derja_Dataset.tunisian-derja-unified-raw-corpusTunisian Derja Unified Raw Corpus
Dataset Description
Repository: hamzabouajila/tunisian-derja-unified-raw-corpus
Paper: Not yet published; dataset card serves as primary documentation
Point of Contact: Hamza Bouajila
License: CC-BY-SA-4.0
Dataset Summary
The Tunisian Derja Unified Raw Corpus is a comprehensive collection of ~802,659 text examples in Tunisian Arabic (Derja), a low-resource dialect of Arabic widely spoken in Tunisia. This raw corpus aggregates data from multiple sources… See the full description on the dataset page: https://huggingface.co/datasets/hamzabouajila/tunisian-derja-unified-raw-corpus.Tunisian_Dialectic_English_Derja
Tunisian-English Dialectic Derja Dataset
Overview
This dataset is a rich and extensive collection of Tunisian dialectic (Derja) and English translations from various sources, updated as of October 2024. It includes synthetic translations, instructional data, media transcripts, social media content, and more.
Dataset Structure
The dataset is composed of JSON files, each containing a list of dictionaries with a text field. The data includes translations… See the full description on the dataset page: https://huggingface.co/datasets/khaled123/Tunisian_Dialectic_English_Derja.SFT-Tunisian-DerjaDPO-Tunisian-Derjafw_tunisian_derjafw_tunisian_derja is a translated version of sawalni-ai/fw-darija with wghezaiel/araT5-marrocan-tunisian model.
derja_to_msa_dataset
Dataset Card
MADAR (License): We construct 12,000 translation instructions using the Multi-Arabic Dialect Applications and Resources (MADAR) corpus is a collection of parallel sentences covering the dialects of 25 Arab cities. We select the dialect of Tunis city as Derja, along with MSA resulting into two translation directions.
SFT_derja_datasetTunisian_Dialectic_English_Derja
Tunisian-English Dialectic Derja Dataset
Overview
This dataset is a rich and extensive collection of Tunisian dialectic (Derja) and English translations from various sources, updated as of October 2024. It includes synthetic translations, instructional data, media transcripts, social media content, and more.
Dataset Structure
The dataset is composed of JSON files, each containing a list of dictionaries with a text field. The data includes translations… See the full description on the dataset page: https://huggingface.co/datasets/abdelfetteh/Tunisian_Dialectic_English_Derja.derja-eng-instructderja
For reference on dataset card metadata, see the spec: https://github.com/huggingface/hub-docs/blob/main/datasetcard.md?plain=1
Doc / guide: https://huggingface.co/docs/hub/datasets-cards
{}
Dataset Card for Tunisian Derja to English Translation Dataset
Dataset Details
Dataset Description
This dataset contains approximately 13,036 rows of sentence-level translations from Tunisian Derja (a dialect of Arabic) to English. The… See the full description on the dataset page: https://huggingface.co/datasets/achermiti/derja.
