datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaps_jpn
Dataset Card for "gaps_jpn"
More Information needed
seamless-align-enA-jpnJpnMix
JpnMix (https://arxiv.org/abs/2512.18834) is a Japanese pretraining corpus built by combining five publicly available Japanese datasets, applying Japanese-specific quality filtering, and performing cross-dataset deduplication.
Subsets
Subset
Description
quality_filtered
Quality-filtered data before deduplication
minhash_deduped
Document-level MinHash deduplication
matched
Documents appearing in 2+ source datasets
The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/JpnMix.Tatoeba-Challenge-jpn-kor
Dataset Card for Dataset Name
This dataset contains Japanese-Korean paired text which is from Helsinki-NLP/Tatoeba-Challenge.
Dataset Details
Dataset Sources
Repository: Helsinki-NLP/Tatoeba-Challenge
Detail: Japanese - Korean jpn-kor
Uses
The dataset can be used to train the translation model that translates Japanese sentence to Korean.
Out-of-Scope Use
You cannot use this dataset to train the model which is to be used under commercial… See the full description on the dataset page: https://huggingface.co/datasets/sappho192/Tatoeba-Challenge-jpn-kor.goldfish-jpn-jpan-100mb-tokenizedcmn_jpn_trainSPEEED_s3_words_fr_ger_jpn_spa_multiOpenAI-MRCR-Translation-JPN本データセットは、ロングコンテキス評価データセット OpenAI MRCR を翻訳して作成した、日本語版の MRCR の評価データセットです。
翻訳方法
日本語版の作成にあたっては、元の英語データの構造を維持しつつ、LLM(Qwen3-235B-A22B)を用いて日本語版のテキストを生成しました。
作成方法としては、
会話履歴を分解
パーツごとに翻訳
元の順番に会話を並べなおす
といった手順により、日本語版の評価サンプルを作成しました。
翻訳が必要だったのは主にタスクの説明文、ユーザの問い合わせ文章、ユーザの最終問い合わせ文章の3箇所です。
タスクの説明文については固定のプロンプトなので、プロンプト全体を一度だけ翻訳しました。
ユーザの問い合わせ文については、全てのユーザの問い合わせが「write a (Document-Type) about (Genre)」という形式の英文になっていたため、Document-Type, Genre の位置の語句を抜き出して翻訳し「(Genre) についての (Document-Type)… See the full description on the dataset page: https://huggingface.co/datasets/abeja/OpenAI-MRCR-Translation-JPN.prosocial-dialog-jpn_Jpanmixture_ami_jpn_synthetic_Bigmixture_callhome_jpn_syntheticdnr-v3-jpnCT-RATE-JPN
CT-RATE-JPN Dataset
CT-RATE-JPN is a Japanese-translated version of radiology reports from the CT-RATE dataset, which contains chest CT volumes paired with corresponding radiology reports.
Dataset Overview
CT-RATE-JPN provides Japanese translations of radiology reports from the CT-RATE dataset to facilitate Japanese medical AI model development. While the original CT-RATE contains 25,692 non-contrast chest CT volumes with corresponding reports, this repository focuses on… See the full description on the dataset page: https://huggingface.co/datasets/YYama0/CT-RATE-JPN.synthetic_dataset_jpn_2_more_speakersjpn-bench
JPN-Bench
JPN-Bench is a Japanese literacy benchmark for tokenizer evaluation and future
Japanese LLM evaluation. This public release contains a small curated
tokenizer-literacy dev set plus benchmark-lane source material manifests kept
separate from tokenizer training material.
This dataset is grouped with the KotodamaLM tokenizer work in the Hugging Face
collection "KotodamaLM Japanese Language Infrastructure".
Files
data/literacy_items.jsonl: 60 tokenizer-literacy… See the full description on the dataset page: https://huggingface.co/datasets/MarcoDotIO/jpn-bench.DiffSinger_opencpop_JPNsynthetic_dataset_jpnsynthetic_dataset_jpn_volumeswmt25-jpn-zhonllb_eng_Latn_jpn_Jpan_subset_15ksynthetic_dataset_jpn_2MentalChat16K
🗣️ Synthetic Counseling Conversations Dataset
📝 Description
Synthetic Data 10K
This dataset consists of 9,775 synthetic conversations between a counselor and a client, covering 33 mental health topics such as 💑 Relationships, 😟 Anxiety, 😔 Depression, 🤗 Intimacy, and 👨👩👧👦 Family Conflict. The conversations were generated using the OpenAI GPT-3.5 Turbo model and a customized adaptation of the Airoboros self-generation framework.
The Airoboros… See the full description on the dataset page: https://huggingface.co/datasets/Jpnm89/MentalChat16K.fineweb-c-jpn-previewgoogle__gemma-2-2b-jpn-it-details
Dataset Card for Evaluation run of google/gemma-2-2b-jpn-it
Dataset automatically created during the evaluation run of model google/gemma-2-2b-jpn-it
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__gemma-2-2b-jpn-it-details.synthetic_dataset_jpn_2_bigmixture_callfriend_jpn_syntheticJP-novel-selected-datasetymcki__gemma-2-2b-jpn-it-abliterated-18-details
Dataset Card for Evaluation run of ymcki/gemma-2-2b-jpn-it-abliterated-18
Dataset automatically created during the evaluation run of model ymcki/gemma-2-2b-jpn-it-abliterated-18
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ymcki__gemma-2-2b-jpn-it-abliterated-18-details.Style_vectors_JPNcpi_jpn
