CoolFace
Datasetpublic

werty1248/sentence-transformer-parallel-En-Ko-with-Similarity

Preprocessing En-Ko subset of Parallel Sentences Datasets 해외에서 제작된 많은 대규모 번역 쌍 데이터들이 영어 텍스트와 한국어 텍스트를 문장 단위로 분리한 후 기계적으로 매핑시키고 있습니다. 이로 인해 전혀 엉뚱한 문장이 번역 쌍으로 매칭되어 있는 문제가 발생합니다. 임베딩 유사도 기반 전처리를 통해 이 문제를 해결할 수 있을 것 같아서 이 데이터 셋을 제작했습니다. 일부 데이터는 상업적 사용이 어려운 라이선스가 적용된 경우가 있습니다. 데이터 목록 parallel-sentences-global-voices parallel-sentences-muse parallel-sentences-jw300 parallel-sentences-tatoeba parallel-sentences-wikititles 유사도 측정 BAAI/BGE-m3로 임베딩 영어… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/sentence-transformer-parallel-En-Ko-with-Similarity.

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes17downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
werty1248/sentence-transformer-parallel-En-Ko-with-Similarity · CoolFace