CoolFace
Datasetpublic

Ik45/wikipedia_dataset_science_en_id

Wikipedia Dataset Science (English - Indonesian) Dataset Description This dataset contains 122,433 aligned sentence pairs extracted from Wikipedia science articles in English and Indonesian. It is highly suitable for Natural Language Processing (NLP) tasks such as machine translation, cross-lingual alignment, and fine-tuning Large Language Models (LLMs) to better understand scientific terminology in Indonesian. Language(s): English (en) and Indonesian (id)… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/wikipedia_dataset_science_en_id.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes35downloads
18 commits on main
6a231496mo ago

Upload folder using huggingface_hub

Ik45
38940046mo ago

Upload folder using huggingface_hub

Ik45
61d806b6mo ago

Update README.md

Ik45
14ca05e6mo ago

Update README.md

Ik45
a63b1236mo ago

Update README.md

Ik45
5bbb93b6mo ago

Update README.md

Ik45
5a458e86mo ago

Rename parallel_dataset_science_en_id.parquet to science_translation/parallel_dataset_science_en_id.parquet

Ik45
41840686mo ago

Upload folder using huggingface_hub

Ik45
e7778ab7mo ago

Upload parallel_dataset_science_en_id.parquet

Ik45
b55df867mo ago

Upload workflow.png

Ik45
db3a7cf7mo ago

Update README.md

Ik45
920ce797mo ago

Update README.md

Ik45
f9bccee7mo ago

Update README.md

Ik45
e4b19d77mo ago

Update README.md

Ik45
59e950f7mo ago

Create README.md

Ik45
a826a777mo ago

Upload crawling_and_cleaning.ipynb

Ik45
d6ceb297mo ago

Upload parallel_dataset_science_en_id.parquet with huggingface_hub

Ik45
a61f1597mo ago

initial commit

Ik45