CoolFace
Datasetpublicgated

InfoBayAI/Indonesian-Non-STEM-Textbook-Dataset

Dataset Description: This dataset is a large-scale collection of Indonesian Non-STEM textbook data, containing 4,098 books and 182.10 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Bahasa. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Indonesian-Non-STEM-Textbook-Dataset.

sourceHugging Facecc-by-4.0updated 8d agoView on Hugging Face
0likes25downloads
fileBasic Concepts and Economic System.parquet188 KBdownload
fileBusiness Communication.parquet138 KBdownload
fileCorporate Law.parquet136 KBdownload
fileCultural Heritage Management Concept for Tourism.parquet167 KBdownload
fileCulture and Identity.parquet167 KBdownload
fileGender, Sexual Health and Reproductive Health Services.parquet149 KBdownload
fileIndustrial Economics.parquet186 KBdownload
fileJuvenile Criminal Law.parquet205 KBdownload
fileMaritime History.parquet201 KBdownload
fileWorkforce Planning.parquet98 KBdownload

InfoBayAI/Indonesian-Non-STEM-Textbook-Dataset · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.