jpn
gemma-2-2b-jpn-itspeaker-segmentation-fine-tuned-callhome-jpnjpn-100mb-after-wc-uniform-oldlex-jpn-ckpt500_seed10_seed10jpn-100mb-after-wc-uniform-oldlex-jpn-ckpt500_seed3407_seed3407jpn-100mb-after-wc-uniform-oldlex-jpn-ckpt500_seed455_seed455gemma-2-2b-jpn-it-i1-GGUFeng-100mb-after-jpn-baseline-newlexicon-zipf-ckpt500_seed3407eng-100mb-after-wc-zipf-newlex-jpn-ckpt500_seed10_seed10
Datasets
All datasets matching “jpn”gaps_jpn
Dataset Card for "gaps_jpn"
More Information needed
seamless-align-enA-jpnJpnMix
JpnMix (https://arxiv.org/abs/2512.18834) is a Japanese pretraining corpus built by combining five publicly available Japanese datasets, applying Japanese-specific quality filtering, and performing cross-dataset deduplication.
Subsets
Subset
Description
quality_filtered
Quality-filtered data before deduplication
minhash_deduped
Document-level MinHash deduplication
matched
Documents appearing in 2+ source datasets
The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/JpnMix.Tatoeba-Challenge-jpn-kor
Dataset Card for Dataset Name
This dataset contains Japanese-Korean paired text which is from Helsinki-NLP/Tatoeba-Challenge.
Dataset Details
Dataset Sources
Repository: Helsinki-NLP/Tatoeba-Challenge
Detail: Japanese - Korean jpn-kor
Uses
The dataset can be used to train the translation model that translates Japanese sentence to Korean.
Out-of-Scope Use
You cannot use this dataset to train the model which is to be used under commercial… See the full description on the dataset page: https://huggingface.co/datasets/sappho192/Tatoeba-Challenge-jpn-kor.goldfish-jpn-jpan-100mb-tokenizedcmn_jpn_train
