CoolFace
Datasetpublic

bekan/english_karakalpak_pairs_parallel_corpus_v2_8907

English-Karakalpak Parallel Corpus v2 (8.9K) Dataset Description English-Karakalpak Parallel Corpus v2 is a high-quality dataset containing 8,906 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. This… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_pairs_parallel_corpus_v2_8907.

sourceHugging Facemitupdated 10mo agoView on Hugging Face
1likes4downloads
11 commits on main
f3fe82710mo ago

Update README.md

bekan
4474ff210mo ago

Upload en_kaa_parallel_corpus_v2_8907.csv

bekan
1ad3a4c10mo ago

Delete en_kaa_v2_8907.csv

bekan
75e168c10mo ago

Update README.md

bekan
e4c352711mo ago

Update README.md

bekan
6b1abf411mo ago

Edit from Data Studio

bekan
0a7003011mo ago

Edit from Data Studio

bekan
708016c11mo ago

Update README.md

bekan
31aed8311mo ago

Update README.md

bekan
d69043511mo ago

Initial commit

bekan
ece698511mo ago

initial commit

bekan