CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01linagora /Tunisian_Derja_Dataset Tunisian Derja Dataset (Tunisian Arabic) This is a collection of Tunisian dialect textual documents for Language Modeling. It was used for continual pre-training of Jais LLM model to Tunisian dialect Languages The dataset is available in Tunisian Arabic (Derja). Data Table subset Lines Tunisian_Dialectic_English_Derja 1220712 HkayetErwi 966 Derja_tunsi 13037 TunBERT 67219 TunSwitchCodeSwitching 394163 TunSwitchTunisiaOnly 380546 Tweet_TN… See the full description on the dataset page: https://huggingface.co/datasets/linagora/Tunisian_Derja_Dataset.text1M<n<10M5 likes210 downloads1y agoHugging Face02hamzabouajila /tunisian-derja-unified-raw-corpusTunisian Derja Unified Raw Corpus Dataset Description Repository: hamzabouajila/tunisian-derja-unified-raw-corpus Paper: Not yet published; dataset card serves as primary documentation Point of Contact: Hamza Bouajila License: CC-BY-SA-4.0 Dataset Summary The Tunisian Derja Unified Raw Corpus is a comprehensive collection of ~802,659 text examples in Tunisian Arabic (Derja), a low-resource dialect of Arabic widely spoken in Tunisia. This raw corpus aggregates data from multiple sources… See the full description on the dataset page: https://huggingface.co/datasets/hamzabouajila/tunisian-derja-unified-raw-corpus.texttext-generation100K<n<1M0 likes108 downloads1y agoHugging Face03khaled123 /Tunisian_Dialectic_English_Derja Tunisian-English Dialectic Derja Dataset Overview This dataset is a rich and extensive collection of Tunisian dialectic (Derja) and English translations from various sources, updated as of October 2024. It includes synthetic translations, instructional data, media transcripts, social media content, and more. Dataset Structure The dataset is composed of JSON files, each containing a list of dictionaries with a text field. The data includes translations… See the full description on the dataset page: https://huggingface.co/datasets/khaled123/Tunisian_Dialectic_English_Derja.text1M<n<10M9 likes106 downloads2y agoHugging Face04wghezaiel /SFT-Tunisian-Derjatext10K<n<100K0 likes22 downloads1y agoHugging Face05wghezaiel /DPO-Tunisian-Derjatext10K<n<100K0 likes21 downloads1y agoHugging Face06wghezaiel /fw_tunisian_derjafw_tunisian_derja is a translated version of sawalni-ai/fw-darija with wghezaiel/araT5-marrocan-tunisian model. text10K<n<100K0 likes17 downloads1y agoHugging Face07wghezaiel /derja_to_msa_dataset Dataset Card MADAR (License): We construct 12,000 translation instructions using the Multi-Arabic Dialect Applications and Resources (MADAR) corpus is a collection of parallel sentences covering the dialects of 25 Arab cities. We select the dialect of Tunis city as Derja, along with MSA resulting into two translation directions. textquestion-answering10K<n<100K0 likes14 downloads1y agoHugging Face08wghezaiel /SFT_derja_datasettext10K<n<100K0 likes13 downloads1y agoHugging Face09abdelfetteh /Tunisian_Dialectic_English_Derja Tunisian-English Dialectic Derja Dataset Overview This dataset is a rich and extensive collection of Tunisian dialectic (Derja) and English translations from various sources, updated as of October 2024. It includes synthetic translations, instructional data, media transcripts, social media content, and more. Dataset Structure The dataset is composed of JSON files, each containing a list of dictionaries with a text field. The data includes translations… See the full description on the dataset page: https://huggingface.co/datasets/abdelfetteh/Tunisian_Dialectic_English_Derja.text1M<n<10M0 likes9 downloads7mo agoHugging Face10hamzabouajila /derja-eng-instructtext10K<n<100K0 likes8 downloads9mo agoHugging Face11achermiti /derja For reference on dataset card metadata, see the spec: https://github.com/huggingface/hub-docs/blob/main/datasetcard.md?plain=1 Doc / guide: https://huggingface.co/docs/hub/datasets-cards {} Dataset Card for Tunisian Derja to English Translation Dataset Dataset Details Dataset Description This dataset contains approximately 13,036 rows of sentence-level translations from Tunisian Derja (a dialect of Arabic) to English. The… See the full description on the dataset page: https://huggingface.co/datasets/achermiti/derja.text10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.