datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
raw-text-corpus
📝 Zomi Raw Text Corpus (Community-Contributed)
The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks.
This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately.
📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.ARABIC-RAW-TEXTDataset:
Aluka 1.4 GB
AraWiki 3.9 GB
Aya 22.5 GB
Islamic Books 21.4 GB
ARABIC-RAW-TEXTDataset:
Aluka 1.4 GB
AraWiki 3.9 GB
Aya 22.5 GB
Islamic Books 21.4 GB
