datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Malaysian-Ultrachat
Ultrachat like using Malaysian context
Prepare multiturn dialogue between user and assistant for malaysian context,
Astroawani, https://huggingface.co/datasets/malaysia-ai/crawl-astroawani, ultrachat-astroawani-malay.jsonl, 60198 rows, 477 MB.
Crossref melayu papers, https://huggingface.co/datasets/mesolitica/crawl-my-website/resolve/main/melayu-pdf.jsonl, ultrachat-crossref-melayu-malay.jsonl, 9959 rows, 187 MB
Epenerbitan… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Ultrachat.Malaysia-textbook-cleaned
Malaysia-textbook-cleaned
Cleaned by Kureiwa.
Cleaned version of Scicom-intl/Malaysia-Textbook, which gathers KSSR and KSSM textbooks in PDF format and converts them to text using Qwen/Qwen3-235B-A22B-Instruct-2507. Covers Bahasa Melayu, Chinese, English, Tamil, and Arabic/Jawi subjects.
Cleaning process
The cleaning pipeline (scripts/clean.py) is pure Python stdlib (csv, re), with DuckDB CLI used only for Parquet I/O. Each page's content is processed as follows:… See the full description on the dataset page: https://huggingface.co/datasets/Kureiwa/Malaysia-textbook-cleaned.Malaysia-Personas
Malaysia-Personas
1,867 synthetic personas of Malaysian voters, in the style of NVIDIA Nemotron-Personas. Each persona is anchored to one real, anonymised record from the published GE15 (2022) electoral roll and fleshed out by an LLM, with calibration to district income and poverty statistics.
They were built for survey simulation: asking a representative synthetic population how it would react to a policy or event, then aggregating the answers.
How it was made… See the full description on the dataset page: https://huggingface.co/datasets/manfye/Malaysia-Personas.Malaysian-Ultrachat
Ultrachat like using Malaysian context
Prepare multiturn dialogue between user and assistant for malaysian context,
Astroawani, https://huggingface.co/datasets/malaysia-ai/crawl-astroawani, ultrachat-astroawani-malay.jsonl, 60198 rows, 477 MB.
Crossref melayu papers, https://huggingface.co/datasets/mesolitica/crawl-my-website/resolve/main/melayu-pdf.jsonl, ultrachat-crossref-melayu-malay.jsonl, 9959 rows, 187 MB
Epenerbitan… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/Malaysian-Ultrachat.
