CoolFace
Datasetpublic

werty1248/sentence-transformer-parallel-En-Ko-with-Similarity

Preprocessing En-Ko subset of Parallel Sentences Datasets 해외에서 제작된 많은 대규모 번역 쌍 데이터들이 영어 텍스트와 한국어 텍스트를 문장 단위로 분리한 후 기계적으로 매핑시키고 있습니다. 이로 인해 전혀 엉뚱한 문장이 번역 쌍으로 매칭되어 있는 문제가 발생합니다. 임베딩 유사도 기반 전처리를 통해 이 문제를 해결할 수 있을 것 같아서 이 데이터 셋을 제작했습니다. 일부 데이터는 상업적 사용이 어려운 라이선스가 적용된 경우가 있습니다. 데이터 목록 parallel-sentences-global-voices parallel-sentences-muse parallel-sentences-jw300 parallel-sentences-tatoeba parallel-sentences-wikititles 유사도 측정 BAAI/BGE-m3로 임베딩 영어… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/sentence-transformer-parallel-En-Ko-with-Similarity.

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes17downloads
settings

This repository belongs to werty1248 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namesentence-transformer-parallel-En-Ko-with-Similarity
visibilitypublic
licencenot set
gatedno
ownerwerty1248
Account settings
werty1248/sentence-transformer-parallel-En-Ko-with-Similarity · CoolFace