datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indonesian-children-books
Indonesian Children Books Dataset
This dataset contains text extracted from Indonesian children's books. This dataset is still contains raw text directly extracted from books, therefore still considered as dirty and need to be preprocessed further.
Dataset Description
Dataset Statistics
Number of books: 2,740
Total pages: 165,245
Total number of tokens: 25,759,439
Number of unique tokens: 698,094
Average tokens per page: 155.89
Extraction Methods… See the full description on the dataset page: https://huggingface.co/datasets/haznitrama/indonesian-children-books.gutenberg_children_books_cleanedgutenberg_children_books_rawMikiV-gutenberg_children_books_cleaned-chunked-8192
