carry08/bo_zh_translation
Dataset source From CUTE (Chinese, Uyghur, Tibetan, English), a large-scale multilingual dataset. Extract paragraphs from parallel corpus, match based on embedding vector stores and semantic search, split into 52381 samples. Usage In alpaca format (Alpaca: A Strong, Replicable Instruction-Following Model), translation bo-zh sentence pairs, can be used in SFT.alpaca_small.json 500 samples sentence average_character_length input 77.97 output 20.93… See the full description on the dataset page: https://huggingface.co/datasets/carry08/bo_zh_translation.
024
Upload 2 files
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Duplicate from CMLI-NLP/CUTE-Datasets
