MarcoDotIO/jpn-bench
JPN-Bench JPN-Bench is a Japanese literacy benchmark for tokenizer evaluation and future Japanese LLM evaluation. This public release contains a small curated tokenizer-literacy dev set plus benchmark-lane source material manifests kept separate from tokenizer training material. This dataset is grouped with the KotodamaLM tokenizer work in the Hugging Face collection "KotodamaLM Japanese Language Infrastructure". Files data/literacy_items.jsonl: 60… See the full description on the dataset page: https://huggingface.co/datasets/MarcoDotIO/jpn-bench.
039
Publish JPN-Bench dataset export
Publish JPN-Bench dataset export
initial commit
