carry08/bo_zh_translation
Dataset source From CUTE (Chinese, Uyghur, Tibetan, English), a large-scale multilingual dataset. Extract paragraphs from parallel corpus, match based on embedding vector stores and semantic search, split into 52381 samples. Usage In alpaca format (Alpaca: A Strong, Replicable Instruction-Following Model), translation bo-zh sentence pairs, can be used in SFT.alpaca_small.json 500 samples sentence average_character_length input 77.97 output 20.93… See the full description on the dataset page: https://huggingface.co/datasets/carry08/bo_zh_translation.
Dataset source
From CUTE (Chinese, Uyghur, Tibetan, English), a large-scale multilingual dataset. Extract paragraphs from parallel corpus, match based on embedding vector stores and semantic search, split into 52381 samples.
Usage
In alpaca format (Alpaca: A Strong, Replicable Instruction-Following Model), translation bo-zh sentence pairs, can be used in SFT. alpaca_small.json 500 samples
alpacamixed.json 52381 samples | sentence | averagecharacter_length | |----------|--------------------------| | input | 108.48 | | output | 38.25 |
