dogdoh/dari-enko-corpus-sample
Dari EN↔KO Technical Translation Corpus (Sample) 🌉 Dari (다리, "bridge") — 67M+ EN↔KO parallel sentence pairs for technical translation. Dataset Description This is a 10,000-pair sample from the full Dari corpus. The full corpus contains 67.2 million high-quality EN↔KO parallel segments across multiple technical domains. Domains Domain Full Corpus Pairs Patent (KIPRIS) 15M+ Medical (PubMed) 12M+ IT/Software 10M+ Legal 8M+ General… See the full description on the dataset page: https://huggingface.co/datasets/dogdoh/dari-enko-corpus-sample.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face