Ik45/wikipedia_dataset_science_en_id
Wikipedia Dataset Science (English - Indonesian) Dataset Description This dataset contains 122,433 aligned sentence pairs extracted from Wikipedia science articles in English and Indonesian. It is highly suitable for Natural Language Processing (NLP) tasks such as machine translation, cross-lingual alignment, and fine-tuning Large Language Models (LLMs) to better understand scientific terminology in Indonesian. Language(s): English (en) and Indonesian (id)… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/wikipedia_dataset_science_en_id.
035
