datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian_Dialectic_English_Derja
Tunisian-English Dialectic Derja Dataset
Overview
This dataset is a rich and extensive collection of Tunisian dialectic (Derja) and English translations from various sources, updated as of October 2024. It includes synthetic translations, instructional data, media transcripts, social media content, and more.
Dataset Structure
The dataset is composed of JSON files, each containing a list of dictionaries with a text field. The data includes translations… See the full description on the dataset page: https://huggingface.co/datasets/khaled123/Tunisian_Dialectic_English_Derja.Tunisian_Dialect_Corpus
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/arbml/Tunisian_Dialect_Corpus.tunisian-dialect-corpus
Dataset Card for Tunisian Dialect Corpus (Cleaned Arabic-Only)
1. Dataset Overview
1.1 Description
This dataset is a cleaned corpus of Tunisian Arabic dialect text, aggregated from multiple public sources on Hugging Face. It is designed for Continual Pretraining (CPT) and general NLP research.
A dedicated preprocessing pipeline was applied to:
Normalize text
Remove noise and artifacts
Filter non-Arabic content
Ensure higher overall data quality
The final… See the full description on the dataset page: https://huggingface.co/datasets/Syrinesmati/tunisian-dialect-corpus.tunisian_dialectData for tunisian dialect corpus: scraped from various forums and aggregated with existing data from Facebook and Twitter.
---
num_examples: 288691
- split: train
Tunisian_Dialectic_English_Derja
Tunisian-English Dialectic Derja Dataset
Overview
This dataset is a rich and extensive collection of Tunisian dialectic (Derja) and English translations from various sources, updated as of October 2024. It includes synthetic translations, instructional data, media transcripts, social media content, and more.
Dataset Structure
The dataset is composed of JSON files, each containing a list of dictionaries with a text field. The data includes translations… See the full description on the dataset page: https://huggingface.co/datasets/abdelfetteh/Tunisian_Dialectic_English_Derja.
