Ik45/wikipedia_dataset_science_en_id
Wikipedia Dataset Science (English - Indonesian) Dataset Description This dataset contains 122,433 aligned sentence pairs extracted from Wikipedia science articles in English and Indonesian. It is highly suitable for Natural Language Processing (NLP) tasks such as machine translation, cross-lingual alignment, and fine-tuning Large Language Models (LLMs) to better understand scientific terminology in Indonesian. Language(s): English (en) and Indonesian (id)… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/wikipedia_dataset_science_en_id.
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Update README.md
Update README.md
Update README.md
Update README.md
Rename parallel_dataset_science_en_id.parquet to science_translation/parallel_dataset_science_en_id.parquet
Upload folder using huggingface_hub
Upload parallel_dataset_science_en_id.parquet
Upload workflow.png
Update README.md
Update README.md
Update README.md
Update README.md
Create README.md
Upload crawling_and_cleaning.ipynb
Upload parallel_dataset_science_en_id.parquet with huggingface_hub
initial commit
