CoolFace
Datasetpublic

yumatin/BiST

BiST 💻 Github Repo English | 简体中文 Introduction BiST is a large-scale bilingual translation dataset, with "BiST" standing for Bilingual Synthetic Translation dataset. Currently, the dataset contains approximately 60M entries and will continue to expand in the future. BiST consists of two subsets, namely en-zh and zh-en, where the former represents the source language, collected from public data as real-world content; the latter represents the target… See the full description on the dataset page: https://huggingface.co/datasets/yumatin/BiST.

sourceHugging Facecc-by-sa-4.0updated 3mo agoView on Hugging Face
0likes110downloads
1 commits on main
d23088c3mo ago

Duplicate from Mxode/BiST

yumatin, Mxode