CoolFace
Datasetpublic

carry08/bo_zh_translation

Dataset source From CUTE (Chinese, Uyghur, Tibetan, English), a large-scale multilingual dataset. Extract paragraphs from parallel corpus, match based on embedding vector stores and semantic search, split into 52381 samples. Usage In alpaca format (Alpaca: A Strong, Replicable Instruction-Following Model), translation bo-zh sentence pairs, can be used in SFT.alpaca_small.json 500 samples sentence average_character_length input 77.97 output 20.93… See the full description on the dataset page: https://huggingface.co/datasets/carry08/bo_zh_translation.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes24downloads
7 commits on main
a46f47f3mo ago

Upload 2 files

carry08
28005c73mo ago

Update README.md

carry08
5ab2a213mo ago

Update README.md

carry08
53be6693mo ago

Update README.md

carry08
01e87413mo ago

Update README.md

carry08
5c5322c3mo ago

Update README.md

carry08
1f9ddef3mo ago

Duplicate from CMLI-NLP/CUTE-Datasets

carry08, YoLo2000