datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leaderboard-requestsllm-jp-corpus-v4-ja_wiki
llm-jp-corpus-v4 — ja_wiki
Mirror of the ja/ja_wiki sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_wiki
Files: 6 × jsonl.gz (1.9 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY-SA 3.0 — inherited from the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_wiki.scaling-data-constrained-llms
Scaling Data-Constrained Language Models with Synthetic Data
This repository provides the pre-training corpora used in Scaling Data-Constrained Language Models with Synthetic Data (Findings of EACL 2026).
Overview
This repository contains multiple corpora designed to study data augmentation strategies for pre-training Japanese LLMs under a data-constrained data setting.
Starting from a limited Japanese Web corpus and a larger English Web corpus, we construct three… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/scaling-data-constrained-llms.llm-jp-corpus-v4-ja_warp_pdf
llm-jp-corpus-v4 — ja_warp_pdf
Mirror of the ja/ja_warp_pdf sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_warp_pdf
Files: 513 × jsonl.gz (73.8 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_pdf.llm-jp-corpus-v4-ja_sip_comprehensive_html
llm-jp-corpus-v4 — ja_sip_comprehensive_html
Mirror of the ja/ja_sip_comprehensive_html sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_html
Files: 181 × jsonl.gz (23.4 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_html.llm-jp-corpus-v4-ja_sip_comprehensive_pdf
llm-jp-corpus-v4 — ja_sip_comprehensive_pdf
Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_pdf
Files: 156 × jsonl.gz (39.1 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.llm-jp-4-thinking-sft-data-chatmlllm-jpのデータセットllm-jp-4-thinking-sft-dataを、
ChatML形式に変換したものです。
ライセンス
各サンプルのライセンスは、元データセットカードに記載された各データソースのライセンスに従います。
本リポジトリは、元となったデータ全体に対して新たなライセンスを付与するものではありません。
利用する場合は、対応する元データソースのライセンス条件を確認してください。
llm-jp-corpus-v4-ja_patent
llm-jp-corpus-v4 — ja_patent
Mirror of the ja/ja_patent sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_patent
Files: 621 × jsonl.gz (58.2 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 — inherited from… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_patent.magpie-sft-v1.0
magpie-sft-v1.0
This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This is a dataset of instruction and response pairs created using the Magpie method.
cyberagent/calm3-22b-chat was used for generating the instructions, and Qwen/Qwen2.5-32B-Instruct was used for generating the responses.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/magpie-sft-v1.0.llm-jp-corpus-v4-ja_kokkai_giji
llm-jp-corpus-v4 — ja_kokkai_giji
Mirror of the ja/ja_kokkai_giji sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_kokkai_giji
Files: 12 × jsonl.gz (1.3 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 —… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_kokkai_giji.oasst2-33k-ja
oasst2-33k-ja
This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
The dataset comprises a Japanese translation of an English subset from oasst2, translated using DeepL.
The English subset can be found here.
For the creation of this dataset, we processed data from kunishou/oasst2-135k-ja.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/oasst2-33k-ja.llm-jp-corpus-v4-ja_warp_html
llm-jp-corpus-v4 — ja_warp_html
Mirror of the ja/ja_warp_html sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_warp_html
Files: 45 × jsonl.gz (1.6 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_html.mbpp-ja
mbpp-ja
This repository provides a mbpp dataset translated from English into Japanese by LLM-jp, a collaborative project launched in Japan.
For English to Japanese translation, DeepL was used.
The links of the original mbpp dataset are here(HuggingFace) or here(GitHub).
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.
Namgi Han, Masatoshi Otake, Shintaro Ozaki, Yusuke Miyao.
oasst1-21k-ja
oasst1-21k-ja
This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This dataset is a Japanese translation of an English subset of oasst1 using DeepL.
English subset is here.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.
Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/oasst1-21k-ja.hh-rlhf-12k-ja
hh-rlhf-12k-ja
This repository provides a human preference dataset developed by LLM-jp, a collaborative project launched in Japan.
This dataset is a Japanese translation of a subset of hh-rlhf using DeepL.
This dataset consists of 12,000 entries randomly sampled from hh-rlhf. Specifically, it includes a random selection of 3,000 entries from the training splits of the four groups: harmless-base, helpful-base, helpful-online, and helpful-rejection-sampled. For more information on… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/hh-rlhf-12k-ja.databricks-dolly-15k-ja
databricks-dolly-15k-ja
This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This dataset is a Japanese translation of databricks-dolly-15k using DeepL.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.
Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi Nakamura, Takashi Kodama, Takumi… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/databricks-dolly-15k-ja.llm-jp-corpus-v4-ja_kaken
llm-jp-corpus-v4 — ja_kaken
Mirror of the ja/ja_kaken sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_kaken
Files: 5 × jsonl.gz (1.1 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 — inherited from the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_kaken.llm-jp-corpus-v4-ja_e-gov
llm-jp-corpus-v4 — ja_e-gov
Mirror of the ja/ja_e-gov sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_e-gov
Files: 2 × jsonl.gz (0.1 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 — inherited from the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_e-gov.LLM-jp-Toxicity-Dataset
LLM-jp Toxicity Dataset
日本語有害文書データセット「LLM-jp Toxicity Dataset」
See https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-toxicity-dataset
acro-yamalex-llmjp-4-math-tir
acro-yamalex-llmjp-4-math-tir(データセット)
日本語数学推論のためのTool-Integrated Reasoning (TIR) データセットです。
自然言語による推論とPythonコード実行を組み合わせたマルチターン形式のデータセットで、OpenWebMathから抽出・生成した問題に対してDeepSeek V3を用いてTIR形式の解法を生成しました。
本データセットはFT-LLM2026コンペティションにおける我々のアプローチの一部として構築されました。
データセット概要
項目
値
データ件数
134,834件
言語
日本語
ソース
OpenWebMath
生成モデル
DeepSeek V3 (deepseek-chat)
フォーマット
JSONL(マルチターン対話形式)
データ作成手法
OpenMathReasoning(NVIDIAのAIMO-2優勝手法)の枠組みに基づき、以下の手順で作成しました。… See the full description on the dataset page: https://huggingface.co/datasets/AcroYAMALEX/acro-yamalex-llmjp-4-math-tir.JSFactCheckBench
JSFactCheckBench
概要
JSFactCheckBenchは、LLM応答のファクトチェックを評価するための日本語ベンチマークです。
日本のソーシャルメディア上で流通した誤情報に由来する85問の質問(JSocialFactより)に対して、11 種類の LLM が生成した応答 170 件をベースに、1,933 件の言明を作成しました。
各言明には、(1) 事実性ラベル(Supported / Refuted / NotEnoughInfo、補助ラベル Unevaluable / Unclear)と、(2) 質問に対する役割を表す重要度ラベル(Main / Sub / Irrelevant)の2軸を人手で付与しています。
事実性は、根拠となるWebページを収集した時点を基準時刻として、その根拠に基づいて判定しています。
詳細は、後日公開予定の論文をご覧ください。
Overview
JSFactCheckBench is a Japanese benchmark for evaluating the fact-checking of LLM… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/JSFactCheckBench.acro-yamalex-llmjp-4-math-cot
acro-yamalex-llmjp-4-math-cot(データセット)
日本語数学推論のためのChain-of-Thought (CoT) データセットです。
StackMathQAの問題に対して、DeepSeek V3を用いて日本語CoT形式の解法を生成しました。
本データセットはFT-LLM2026コンペティションにおける我々のアプローチの一部として構築されました。
データセット概要
項目
値
データ件数
306,366件
言語
日本語
ソース
StackMathQA
生成モデル
DeepSeek V3 (deepseek-chat)
フォーマット
JSONL(マルチターン対話形式)
データ作成手法
OpenMathReasoning(NVIDIAのAIMO-2優勝手法)の枠組みに基づき、以下の手順で作成しました。
Step 1: 日本語CoT解法の生成
StackMathQAの問題に対してDeepSeek… See the full description on the dataset page: https://huggingface.co/datasets/AcroYAMALEX/acro-yamalex-llmjp-4-math-cot.oasst1-21k-en
oasst1-21k-en
This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This dataset is an English subset of oasst1.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.
Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi Nakamura, Takashi Kodama, Takumi Okamoto.
llava-instruct-ja
Dataset Card for llava_instruct_ja
Dataset details
This is the Japanese version of LLaVA-Instruct, which contains 156K samples.
We used gpt-4o-mini-2024-07-18 to generate data through via Azure OpenAI API.
License
Creative Commons Attribution 4.0 License; and it should abide by the OpenAI terms of use
llm-jp-corpus-v4-ja_wiki-manufacturing
ja_wiki 製造業コーパス
llm-jp-corpus-v4 の
ja/ja_wiki(120 万記事)から、製造業に関連する文書を抽出した継続事前学習(CPT)用データ。
文書数
47,914 件(元コーパスの 3.99%)
文字数
105M 字(推定 74M トークン)
形式
JSONL・1 行 1 文書
抽出方法
キーワード一致では役に立たない。「製造」「工場」で本文先頭を引くと 5.7% が命中するが、
中身は工場や製造に一言触れただけの記事が大半だった。そこで 精度を Wikipedia の
カテゴリ木、網羅性を分類器に担わせ、その和集合を採っている。
カテゴリ木 — 製造業・業種別・機械工学・材料工学・産業遺産など 42 のシードから
日本語 Wikipedia のカテゴリグラフを深さ 3 まで辿る(SQL ダンプからオフライン構築)。
突合 — 得られたタイトルを ja_wiki の実在記事に絞る。
分類器 — 上記を正例、無作為抽出を負例として、文字… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_wiki-manufacturing.aya-ja-evol-inst
aya-ja-eval-inst
This repository provides a preference dataset developed by LLM-jp, a collaborative project launched in Japan.
This dataset was created by generating the chosen response using Qwen/Qwen2.5-32B-Instruct and the rejected response using llm-jp/llm-jp-3-1.8b-instruct for the prompts in weblab-GENIAC/aya-ja-evol-instruct-calm3-dpo-masked.
This repository does not contain prompts but only the corresponding indices. Please obtain the original prompts from the original data… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/aya-ja-evol-inst.ja-vg-vqa-conversation
Dataset Card for ja-vg-vqa-conversation
Dataset details
This dataset consists of multiple QA pairs from Japanese Visual Genome VQA dataset that have been converted into multi-turn conversations.
The prompt 語句または短い文で答えてください。 has been appended to the questions in the first turn.
We excluded the samples from the JA-VG-VQA-500 dataset, resulting in a total of 98,708 samples.
License
Creative Commons Attribution 4.0 License
japanese-photos-conversation
Dataset Card for japanese photos conversation
Dataset details
This dataset contains multi-turn conversational instructions about images taken in Japan. The images were sourced from https://huggingface.co/datasets/ThePioneer/japanese-photos.
We input each image into GPT-4o (gpt-4o-2024-05-13) via the Azure OpenAI API to generate the instruction data.
Some of the images in the original dataset were filtered by the Azure OpenAI API when they were input, resulting in a total… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/japanese-photos-conversation.jpacifico__Chocolatine-14B-Instruct-4k-DPO-details
Dataset Card for Evaluation run of jpacifico/Chocolatine-14B-Instruct-4k-DPO
Dataset automatically created during the evaluation run of model jpacifico/Chocolatine-14B-Instruct-4k-DPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jpacifico__Chocolatine-14B-Instruct-4k-DPO-details.magpie-llm-jp-3-13b-20k
合成日本語指示データセット
概要
このデータセットは、大規模言語モデル(LLM)を用いて自動生成された日本語の指示とそれに対する応答のコレクションです。データセットは指示付与型のタスクのための学習や評価に使用することを目的としています。
データセット仕様
サンプル数: 20,000
言語: 日本語
フォーマット: JSON
生成方法
データセットは以下のプロセスを通じて生成されました:
LLM-JP 3.13B Instructモデルを使用
各サンプルは3段階のプロセスで生成:
a) 指示文の生成
b) Chain-of-Thought (CoT) 応答の生成 (一部のデータには含まれない)
c) 最終的な応答のself-refine
https://github.com/DeL-TaiseiOzaki/magpie-llm-jp-3
データ構造
各サンプルは以下の構造を持つJSONオブジェクトです:
{
"instruction": "指示文"… See the full description on the dataset page: https://huggingface.co/datasets/DeL-TaiseiOzaki/magpie-llm-jp-3-13b-20k.
