datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JpnMix
JpnMix (https://arxiv.org/abs/2512.18834) is a Japanese pretraining corpus built by combining five publicly available Japanese datasets, applying Japanese-specific quality filtering, and performing cross-dataset deduplication.
Subsets
Subset
Description
quality_filtered
Quality-filtered data before deduplication
minhash_deduped
Document-level MinHash deduplication
matched
Documents appearing in 2+ source datasets
The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/JpnMix.jpn-bench
JPN-Bench
JPN-Bench is a Japanese literacy benchmark for tokenizer evaluation and future
Japanese LLM evaluation. This public release contains a small curated
tokenizer-literacy dev set plus benchmark-lane source material manifests kept
separate from tokenizer training material.
This dataset is grouped with the KotodamaLM tokenizer work in the Hugging Face
collection "KotodamaLM Japanese Language Infrastructure".
Files
data/literacy_items.jsonl: 60 tokenizer-literacy… See the full description on the dataset page: https://huggingface.co/datasets/MarcoDotIO/jpn-bench.babylm-jpn
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: jpn
Script: Hira, Jpan, Kana
Tier: 10M
Byte Premium Factor: 1.321974
Size (MB): 71.78
Expected Size (MB): 71.78
Number of Documents: 2,043
Total Tokens: 16,524,324
Tokenizer: tohoku-nlp/bert-base-japanese
Tokens Per Category
child-books: 9,712,521 tokens
educational: 291,053… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-jpn.
