datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cherokee-english-translation
Cherokee–English Parallel Corpus (Archivist Project)
A curated Cherokee (ᏣᎳᎩ / Tsalagi) ↔ English parallel corpus for machine
translation, assembled from public sources, deduplicated, benchmark-decontaminated,
and conflict-cleaned. Built to train and evaluate English→Cherokee translation
models for one of the most endangered languages in North America.
Files
File
Rows
Purpose
train_en2chr_v2.jsonl
138,307
Flagship training set. English→Cherokee SFT… See the full description on the dataset page: https://huggingface.co/datasets/CGICAI/cherokee-english-translation.cherokee-english-word-10.2k
Cherokee-English Word Dataset (10k)
Overview
The Cherokee-English Word Dataset is a comprehensive collection of 10,000 entries, each containing a word from the Cherokee language along with its English translation. This dataset is designed to facilitate linguistic research, aid in the development of machine translation models, and support educational initiatives aimed at preserving and promoting the Cherokee language.
Data Structure
Each entry in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/wang4067/cherokee-english-word-10.2k.cherokee-english-bible-7.96k
Cherokee-English Bible Dataset (8k)
Overview
The Cherokee-English Bible Dataset is a specialized collection of 8,000 entries, each containing a verse from the Cherokee Bible along with its English translation. This dataset is a valuable resource for linguistic scholars, theologians, and developers working on language processing tools that require a deep understanding of both the Cherokee and English languages, particularly within the context of religious texts.… See the full description on the dataset page: https://huggingface.co/datasets/wang4067/cherokee-english-bible-7.96k.cherokee-english-45
