datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ar-En-Code-Switching-Textual-Dataset
ArE-CSTD: Arabic-English Code-Switching Textual Dataset
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "ArE-CSTD" dataset, which stands for "Arabic-English Code-Switching Textual Dataset”.
This dataset contains 330K dialectical Arabic-English code-swithing sentences generated by the large language model GPT-4.
TXT Files
There are 6 txt files. 2 files for Modern Standard Arabic(MSA) train and… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Ar-En-Code-Switching-Textual-Dataset.zygai_textual
📘 ZygAI Textual Research Dataset v1.0
The ZygAI Textual Research Dataset is a multilingual, multi-domain collection of Lithuanian and English texts.It was created as part of ZygAI Research (2025–2026) to build culturally-aware AI systems focused on Lithuania, Baltic traditions, history, language, and modern society.
This dataset is fully standardized and ideal for LLM fine-tuning, RAG knowledge bases, topic modeling, translation tasks, and historical/cultural analysis.… See the full description on the dataset page: https://huggingface.co/datasets/ZygAI/zygai_textual.
