datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kakugo-mlt
Kakugo Maltese dataset
[Paper] [Code] [Model]
A synthetically generated conversation dataset for training in Maltese.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Maltese. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/kakugo-mlt.yaruo-mlt-aa
Japanese AA (Shift_JIS art) — MLT収集所 collection v32.0
1,381,069 individual AA pieces parsed from the 15,002 MLT compilation files of the
やる夫スレ用MLT収集所
(Yaruo-thread MLT Collection), the community-maintained archive of 2channel/5channel
ASCII art. Obtained from the collection's official まとめzip
(mlt_v32_0.zip, 2026-06-27, distributed via AAHub /
AAMZ Viewer).
This is, to our knowledge, the largest machine-readable corpus of human-made text art
ever assembled: ~25 years of Japanese… See the full description on the dataset page: https://huggingface.co/datasets/1anon8anon1/yaruo-mlt-aa.
