CoolFace
Datasetpublic

Ik45/wikipedia_dataset_science_en_id

Wikipedia Dataset Science (English - Indonesian) Dataset Description This dataset contains 122,433 aligned sentence pairs extracted from Wikipedia science articles in English and Indonesian. It is highly suitable for Natural Language Processing (NLP) tasks such as machine translation, cross-lingual alignment, and fine-tuning Large Language Models (LLMs) to better understand scientific terminology in Indonesian. Language(s): English (en) and Indonesian (id)… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/wikipedia_dataset_science_en_id.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes36downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Ik45/wikipedia_dataset_science_en_id · CoolFace