CoolFace
Datasetpublic

carry08/bo_zh_translation

Dataset source From CUTE (Chinese, Uyghur, Tibetan, English), a large-scale multilingual dataset. Extract paragraphs from parallel corpus, match based on embedding vector stores and semantic search, split into 52381 samples. Usage In alpaca format (Alpaca: A Strong, Replicable Instruction-Following Model), translation bo-zh sentence pairs, can be used in SFT.alpaca_small.json 500 samples sentence average_character_length input 77.97 output 20.93… See the full description on the dataset page: https://huggingface.co/datasets/carry08/bo_zh_translation.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes24downloads
Dataset Card

Dataset source

From CUTE (Chinese, Uyghur, Tibetan, English), a large-scale multilingual dataset. Extract paragraphs from parallel corpus, match based on embedding vector stores and semantic search, split into 52381 samples.

Usage

In alpaca format (Alpaca: A Strong, Replicable Instruction-Following Model), translation bo-zh sentence pairs, can be used in SFT. alpaca_small.json 500 samples

sentenceaverage_character_length
input77.97
output20.93

alpacamixed.json 52381 samples | sentence | averagecharacter_length | |----------|--------------------------| | input | 108.48 | | output | 38.25 |