NNEngine/English-Hindi_Translation
📘 README.md 👉 Copy everything below into your repository README.md English–Hindi Massive Synthetic Translation Dataset 🧠 Overview This dataset is a large-scale synthetic parallel corpus for English → Hindi machine translation, designed to stress-test modern sequence-to-sequence models, tokenizers, and large-scale training pipelines. The corpus contains 10 million aligned sentence pairs generated using a high-entropy template engine with: 100+… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/English-Hindi_Translation.
076
1version https://git-lfs.github.com/spec/v12oid sha256:b3076f23ab1f870c6c451ade36373bd602e7aaa7358ff5dc30cce1777d792f273size 44490114944 