CoolFace
Datasetpublic

afrizalha/Indo4B-Combined

This is the entire Indo4B dataset, combined into a single file. The original dataset can be found here: https://github.com/IndoNLP/indonlu This is a combination of all the different files in the compressed .tar.xz. The goal is so that anyone who's interested in Indonesian NLP can fairly simply load this dataset from huggingface, already combined in full. Note the original files consists of line-separated strings. This dataset just combines them while removing the available blank lines.

sourceHugging Facemitupdated 3y agoView on Hugging Face
0likes569downloads
6 commits on main
0890aff3y ago

Librarian Bot: Add language metadata for dataset (#2)

Catinthebag, librarian-bot
fa5451c3y ago

Update README.md

Catinthebag
c0c96ec3y ago

Update README.md

Catinthebag
614244e3y ago

Upload dataset (part 00001-of-00002)

Catinthebag
e0ba2f23y ago

Upload dataset (part 00000-of-00002)

Catinthebag
c250fab3y ago

initial commit

Catinthebag