traintogpb/aihub-koen-translation-integrated-base-1m
AI Hub Ko-En Translation Dataset (Integrated) AI Hub의 한-영 번역 관련 데이터셋 8개를 병합한 자료입니다. 병합 시 총 데이터 개수는 10,416,509개 이며, train / validation / test는 8:1:1 비율로 분할되었습니다. base-10m: 병합 데이터 100% 사용, 총 10,416,509개 mini-1m: 병합 데이터 10% 사용 (base-10m의 각 세트 내에서 10% 임의 선택), 총 1,041,651개 tiny-100k: 병합 데이터 1% 사용 (base-10m의 각 세트 내에서 1% 임의 선택), 총 104,165개 Subsets 활용한 데이터셋 목록은 다음과 같으며, 데이터셋 이름 옆 번호는 aihubshell에서의 datasetkey입니다. 전문분야 한영 말뭉치 (111) 총 개수: 1,350,000 중복 제거 후 개수: 1,350,000… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-koen-translation-integrated-base-1m.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face