datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-2-edu-japanese
🍷 FineWeb2 Edu Japanese: High-Quality Educational Japanese Dataset
This dataset consists of 120 million texts (approximately 89.3B tokens) filtered from the 376 million Japanese texts in FineWeb2 that were deemed educational. The following subsets are also provided:
default: Approximately 120M texts (120 million texts) totaling around 89.3B tokens
sample_10BT: A random sample of about 10B tokens from the default dataset
small_tokens: Data composed solely of texts with 512 tokens… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-2-edu-japanese.Japanese_NicoNico_Douga_Movie_Meta_Data_2016Japanese-Novels-23M
Japanese-Novels-23M
This dataset contains Japanese web novels that I collected personally.
Machine-Learning Use OnlyAccess is restricted to bona fide machine-learning–related purposes.To request access, please provide a detailed explanation of the specific tasks or applications for which you intend to use the dataset.
Total records: 23,212,809
Total characters: 80,846,120,027
Total tokens (Llama 4 tokenizer): 55,406,468,406 (55.4 B)
laion2B-japanese-subsetJapanese_Bandori_Band_Story
Japanese Bandori Band Story
Japanese Band Story text retrieved from the Bestdori scenario assets.
This snapshot contains 26 story entries, 493 chapters,
and 30679 rows (28800 dialogue rows).
Created at 2026-09-15T02:11:27.707570+00:00.
Files
data/train-*.parquet: Hub dataset shards generated by Dataset.push_to_hub.
data/band_stories.jsonl: local combined dataset, also included in the downloadable ZIP.
stories/story_XXXX/: complete per-story TXT, CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/KomeijiForce/Japanese_Bandori_Band_Story.japanese-conversion
Awesome Japanese IME Training Data
Awesome Japanese Corpus の本文を直接 KyTea で解析し、文脈付きかな漢字変換の
ランキング学習例を作成したデータセットです。中間の読み付きデータセットは
作りません。任意の検証モードでは、抽出範囲についてMeCabの読みとも一致した
例だけを採用できます。
context: 変換対象より前の本文
input: 変換対象のひらがな読み
correct: 元コーパスにある正解表記
incorrect: predict.py で全体または一部分を再変換した誤候補の配列
n_words: 抽出した連続形態素数
source_text と target_start / target_end により、元文章中の抽出位置を
復元できます。元データの利用条件は from と from_license を参照して
ください。
relaion2B-en-research-safe-japanese-translation
relaion2B-en-research-safe-japanese-translation
This dataset is the Japanese translation of the English subset of ReLAION-5B (laion/relaion2B-en-research-safe), translated by gemma-2-9b-it.
We used text2dataset for translating with open-weight LLMs.
By leveraging the fast LLM inference library vLLM, this tool enables the rapid translation of large English datasets into Japanese.
Prompt
The following is the prompt used for translation with Gemma.
You are an… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/relaion2B-en-research-safe-japanese-translation.Japanese-PD
🇯🇵 Japanese Public Domain 🇯🇵
Japanese-Public Domain or Japanese-PD is a large collection aiming to aggregate all Japanese monographies and periodicals in the public domain.
Dataset summary
The collection contains 1,410 titles making up 21,072,188 words recovered from the Internet Archive. Each parquet file has the full text of 2,000 books selected at random.
Curation method
The composition of the dataset adheres to the criteria for public domain works in… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Japanese-PD.fineweb-2-edu-japanese-top30mhotchpotch/fineweb-2-edu-japanese
こちらのデータセットのscoreが3.2以上のデータを抽出したものです。
良質なデータを公開してくださったhotchpotch様に感謝申し上げます。
License
This dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0, as is the original FineWeb2 dataset. Additionally, its use is subject to the CommonCrawl Terms of Use.
fungi_diagnostic_chars_comparison_japanese
fungi_diagnostic_chars_comparison_japanese大菌輪「識別形質まとめ」データセット最終更新日 / Last updated: 2026/8/29(up to R3-14214)
Languages
Japanese
This dataset is available in Japanese only.
概要 / Overview
Atsushi Nakajima(中島淳志)が個人で運営しているWebサイト大菌輪では、数千件以上の菌類分類学論文を「論文3行まとめ」という形で要約および索引付け(インデキシング)した情報を提供しています。その一環として、ある菌と別の菌の「共通する」あるいは「異なる」識別形質 (diagnostic characters) に関する記述を人手で抽出しています。
Daikinrin, a personal website run by Atsushi Nakajima, provides summaries and… See the full description on the dataset page: https://huggingface.co/datasets/Atsushi/fungi_diagnostic_chars_comparison_japanese.MMARCO-japanese-32-scored-triplets@misc{clavié2024jacolbertv25optimisingmultivectorretrievers,
title={JaColBERTv2.5: Optimising Multi-Vector Retrievers to Create State-of-the-Art Japanese Retrievers with Constrained Resources},
author={Benjamin Clavié},
year={2024},
eprint={2407.20750},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2407.20750},
}
japanese-conala
Attribution
MTEB-format derivative of haih2/japanese-conala (Japanese CoNaLa; intent->code). Query = Japanese intent; corpus = Python snippet. Paired test+train combined.
filtered_japanese-wikipediawikipedia-japanese
Japanese Wikipedia Dataset
This dataset is a comprehensive pull of all Japanese wikipedia article data as of 20220808.
Note: Right now its uploaded as a single cleaned gzip file (for faster usage), I'll update this in the future to include a huggingface datasets compatible class and better support for japanese than the existing wikipedia repo.
Example use case:
gunzip jawwiki20200808.json.gz
import pandas as pd
from datasets import load_dataset
df =… See the full description on the dataset page: https://huggingface.co/datasets/inarikami/wikipedia-japanese.japanese-wikipedia-paragraphsA slightly modified version of the parsing and chunking method for singletongue/wikipedia-utils.
Pre-processing was performed using oshizo/wikipedia-utils, which is a fork of the original repository, singletongue/wikipedia-utils.
The Wikipedia data was crawled between 2023/12/5 and 2023/12/8.
aio-passages-bpr-bert-base-japanese-v3
Dataset Card for llm-book/aio-passages-bert-base-japanese-v3-bpr
書籍『大規模言語モデル入門』で使用する、「AI王」コンペティションのパッセージデータセットに BPR によるパッセージの埋め込みを適用したデータセットです。
llm-book/aio-passages のデータセットに対して、llm-book/bert-base-japanese-v3-bpr-passage-encoder によるパッセージのバイナリベクトルが embeddings フィールドに追加されています。
Licence
本データセットで利用している Wikipedia のコンテンツは、クリエイティブ・コモンズ表示・継承ライセンス 3.0 (CC BY-SA 3.0) および GNU 自由文書ライセンス (GFDL) の下に配布されているものです。
fineweb-2-edu-japanese-noise-detect-rawfineweb-2-edu-japanese の small_tokens の text カラムをユニコード正規化(NFKC)したものを fineweb-2-japanese-text-cleaner を使ってノイズ箇所を推論したRAWデータセットです。
このデータセットで、ノイズ文字列を削除したのものを、fineweb-2-edu-japaneseのsmall_tokens_cleanedサブセットとして公開しています。
推論時のパラメータは閾値0.7以上かつノイズの文字列長4文字以上のものを、noise_spans カラムに付与しています。noise_spans は start_pos, end_pos のペアとなってます。
ライセンス
ODC-By
historical-japanese-yen-cross-forwards-sample
Historical Japanese Yen Cross FX Forwards Sample
A free evaluation sample of historical FX-forward data for selected non-USD Japanese yen crosses across multiple tenors.
Full historical FX-forward catalog, broader pair and tenor coverage, downloadable datasets, and pricing:https://futuresforexandsomeindexes.com/
This repository is a free evaluation sample intended for schema inspection, data-quality evaluation, integration testing, and quantitative research prototyping. It is… See the full description on the dataset page: https://huggingface.co/datasets/lynx1231/historical-japanese-yen-cross-forwards-sample.premodern-japanese-books-lm-corpus
Premodern Japanese Books LM Corpus
日本語
概要
日本古典籍統一データセットの言語モデル学習用本文ビュー v0.2.0 です。lm-curated v0.2.0から、本文採用対象とした1,598文書を収録しています。文書の本文はcontent列に入り、文書単位の論理分割はsplit列に記録しています。
収録範囲
kouigenji
ndl-minhon-ocrdataset
yatanavi
利用上の注意
Hugging Face上の物理splitはtrain一つです。split列にtrain、validation、testの論理分割を保持しています。
NDL Minhonでは角括弧の記号だけを削除し、角括弧内部の文字は保持しています。
やたナビでは読み仮名、異読、校訂注、その他の補助表記を除去しています。
利用条件と帰属表示はNOTICE.mdを確認してください。… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/premodern-japanese-books-lm-corpus.jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3
Dataset Card for "jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3"
More Information needed
historical-japanese-yen-crosses-sample
Historical Japanese Yen Crosses Sample
A free evaluation sample of historical Forex data for selected Japanese yen cross-currency pairs.
Full historical Forex catalog, broader pair coverage, downloadable datasets, and pricing:https://futuresforexandsomeindexes.com/
This repository is a free evaluation sample intended for schema inspection, data-quality evaluation, integration testing, and quantitative research prototyping. It is not the complete commercial dataset.… See the full description on the dataset page: https://huggingface.co/datasets/lynx1231/historical-japanese-yen-crosses-sample.premodern-japanese-books-source-crosswalk
Premodern Japanese Books Source Crosswalk
日本語
概要
日本古典籍統一データセット v0.1.0 の公開用Viewです。固定済み内部Releaseから、再配布と機械学習利用が許可された行だけを収録しています。収録行数は 17,570,138 行、収録SourceDataset数は 7 件です。
用途
日本語歴史資料の研究、検索、OCRまたは言語モデル用データ処理に利用できます。個々の行には採用したCanonical ID、権利判定、必要な帰属を保持しています。
権利
単一のライセンス値は全行の条件を表しません。必ず NOTICE.md と行単位の権利列を確認してください。
限界
v0.1.0 は監査対象52候補のうち、固定入力が成立した21 SourceDatasetを対象とする段階公開です。内容の正確性、外部参照の永続性、特定用途への適合性を保証しません。… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/premodern-japanese-books-source-crosswalk.japanese-mt-benchjapanese-stackexchange
japanese-stackexchange
英語による日本語に関する質問ができる Japanese Stack Exchange のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。
日本語翻訳された StackExchange ではないです。
データ構造
投稿本文は html2text を使ってマークダウン化されています。その際、
コードブロックは ``` で囲まれるように変更されています。
画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。
default サブセット
id: 質問投稿の ID
question: 質問投稿
answers: 質問に対する回答投稿のリスト
accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある
popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある
simple サブセット… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/japanese-stackexchange.japanese-qa-reasoning-100k
思考過程を含む、日本語質問・キーワード・回答・文章の合成データセット
fineweb2-edu-japanese の文章データを元に、DeepSeek-R1 で文章(text)から質問文と回答部分の該当箇所を生成した日本語の質問と対応する文章・回答部のデータセットです。deekseek-r1 が出力した reasoning 部分も含まれます。testセットは、fineweb2-edu-japaneseのtestのみからサンプリングしています。
質問と文章ペアのデータセットやキーワードと文章ペアのデータセットとしてお使いいただけます。
ライセンス
fineweb2 と同等の ODC-By とします。
japanese-query-crafter-reasoning-80k
思考過程を含む、クエリ作成のための日本語質問文テキストの合成データセット
fineweb2-edu-japanese の small_tokens_cleaned の文章データを元に、DeepSeek-R1 で文章(text)から質問文を作成したデータセットです。deekseek-r1 が出力した reasoning 部分も含まれます。testセットは、fineweb2-edu-japaneseのtestのみからサンプリングしています。
ライセンス
fineweb2 と同等の ODC-By とします。
japanese-text-difficulty
Aozora Text Difficulty Dataset
This dataset contains Japanese literary texts from the Aozora Bunko digital library, enhanced with jReadability-based difficulty analysis for Japanese language learning and curriculum development.
Dataset Overview
Source: Aozora Bunko (青空文庫) - Japan's premier digital library of public domain literature
Enhancement: jReadability-based difficulty scoring using research-backed Japanese readability models
Primary Methodology: jReadability - A… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-text-difficulty.japanese-math-empirical-difficulty-pilot-50k
Japanese Math Empirical Difficulty Pilot 50k
This dataset is a 50,000-problem empirical difficulty pilot, not a full empirical labeling of the original 5.66M-row source dataset.
It was created for LLM-jp experiment 0399, Team Victory SFT, to validate empirical difficulty label distribution, downstream split behavior, and the rollout/scoring pipeline before attempting labeling at the full 5.6M scale.
Current Status
This upload uses the v3 scorer with assistant-only… See the full description on the dataset page: https://huggingface.co/datasets/argo11/japanese-math-empirical-difficulty-pilot-50k.JParaCrawl-Filtered-English-Japanese-Parallel-Corpus
Introduction
This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus.
The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet.
Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.metatree_JapaneseVowels
Dataset Card for "metatree_JapaneseVowels"
More Information needed
