CoolFace
Datasetpublic

bekan/english_karakalpak_pairs_parallel_corpus_v2_8907

English-Karakalpak Parallel Corpus v2 (8.9K) Dataset Description English-Karakalpak Parallel Corpus v2 is a high-quality dataset containing 8,906 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. This… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_pairs_parallel_corpus_v2_8907.

sourceHugging Facemitupdated 10mo agoView on Hugging Face
1likes4downloads

bekan/english_karakalpak_pairs_parallel_corpus_v2_8907 · main · files are served by the source, never re-hosted here