CoolFace
Datasetpublic

Ik45/data-science-en-id

Data Science EN-ID Parallel Corpus (Scientific Domain) Dataset Description This dataset is a curated English-Indonesian (EN-ID) parallel corpus specifically designed for the Scientific and Data Science domains. It was developed to support the training of Machine Translation (NMT) models and Large Language Models (LLMs) to better handle technical terminology, academic structures, and formal scientific language. Primary Languages: English (EN) and Indonesian (ID)… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/data-science-en-id.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes186downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Ik45/data-science-en-id · CoolFace