CoolFace
20 results

jpn

bjoernp /gaps_jpn Dataset Card for "gaps_jpn" More Information needed text100M<n<1B1 likes530 downloads3y agoHugging Faceasahi417 /seamless-align-enA-jpnaudio100K<n<1M0 likes258 downloads2y agoHugging FaceAdaMLLab /JpnMix JpnMix (https://arxiv.org/abs/2512.18834) is a Japanese pretraining corpus built by combining five publicly available Japanese datasets, applying Japanese-specific quality filtering, and performing cross-dataset deduplication. Subsets Subset Description quality_filtered Quality-filtered data before deduplication minhash_deduped Document-level MinHash deduplication matched Documents appearing in 2+ source datasets The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/JpnMix.texttext-generation100M<n<1B2 likes176 downloads5mo agoHugging Facesappho192 /Tatoeba-Challenge-jpn-kor Dataset Card for Dataset Name This dataset contains Japanese-Korean paired text which is from Helsinki-NLP/Tatoeba-Challenge. Dataset Details Dataset Sources Repository: Helsinki-NLP/Tatoeba-Challenge Detail: Japanese - Korean jpn-kor Uses The dataset can be used to train the translation model that translates Japanese sentence to Korean. Out-of-Scope Use You cannot use this dataset to train the model which is to be used under commercial… See the full description on the dataset page: https://huggingface.co/datasets/sappho192/Tatoeba-Challenge-jpn-kor.texttranslation10M<n<100M0 likes126 downloads3y agoHugging Facefpadovani /goldfish-jpn-jpan-100mb-tokenized100K<n<1M0 likes98 downloads4mo agoHugging Faceyiyic /cmn_jpn_traintext1M<n<10M0 likes94 downloads2y agoHugging Face