datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisper_transcriptions.reazon_speech_allwhisper_transcriptions.reazonspeech.allja_asr.reazon_speech_alljapanese-anime-speech-v2
Japanese Anime Speech Dataset V2
日本語はこちら
japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models.
The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels.
This dataset is not an updated version of japanese-anime-speech-v1.
For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset.
The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2.whisper_transcriptions.reazonspeech.all.wer_10.0en_asr.mlsfineweb-2-edu-japanese
🍷 FineWeb2 Edu Japanese: High-Quality Educational Japanese Dataset
This dataset consists of 120 million texts (approximately 89.3B tokens) filtered from the 376 million Japanese texts in FineWeb2 that were deemed educational. The following subsets are also provided:
default: Approximately 120M texts (120 million texts) totaling around 89.3B tokens
sample_10BT: A random sample of about 10B tokens from the default dataset
small_tokens: Data composed solely of texts with 512 tokens… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-2-edu-japanese.Japanese_NicoNico_Douga_Movie_Meta_Data_2016japanese-anime-speech
Japanese Anime Speech Dataset
日本語はこちら
japanese-anime-speech is an audio-text dataset designed for the training of automatic speech recognition models. The dataset is comprised of thousands of audio clips and their corresponding transcriptions from different visual novels.
The goal of this dataset is to increase the accuracy of automatic speech recognition models, such as OpenAI's Whisper, in accurately transcribing dialogue from anime and other similar Japanese media. This genre is… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/japanese-anime-speech.Japanese-Medical-VQA-12m
Japanese Medical VQA 12M
Japanese Medical VQA 12M is a large-scale Japanese medical multimodal dataset built from Open-PMC-18M and released in Parquet and Webdataset format.
This dataset contains outputs from multiple data-construction stages, including:
source captions
Japanese translations of source captions
enriched captions
Japanese translations of enriched captions
question-answering
Current Repository Format
This repository currently stores the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/MIL-UT/Japanese-Medical-VQA-12m.japanese-text-image-retrieval-trainshunk031/JDocQAのtrain splitに含まれるPDFデータを画像化し、NDLOCRでOCRしたテキストとペアにしたデータセットです。OCRは長い辺を1200pxにリサイズした画像に対して実施しました。OCR結果には、読み取りに失敗した際の文字列「〓」が含まれます。本データセットに含めている画像は、長い辺を896px、700px、588pxのいずれかにリサイズしています。どのサイズとするかは主にページに含まれる文字数で決めました。
query列は、OCR結果の文字列に対しQwen/Qwen2.5-14B-Instructで生成したものです。3つの質問を生成させ、ランダムに1つを選んだものをデータセットに含めました。質問を生成する際は以下のプロンプトを使用しました。
あなたは、質問から画像をretrieveするためのモデルをトレーニングするための(質問, 画像)ペアのデータセットを作成するプロジェクトのメンバーである。
プロジェクトは以下のように進める。
step1. ドキュメントPDFを1ページ1枚の画像ファイルに変換する
step2.… See the full description on the dataset page: https://huggingface.co/datasets/oshizo/japanese-text-image-retrieval-train.Japanese-Creative-Writing-39.6k
Japanese-Creative-Writing-39.6k
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した、約39600件の日本語の小説執筆タスクデータセットです。
全てのデータは2ターンのデータとなっています。また、データセット中の一部データはNSFW表現を含みます。
データの詳細
各データは以下のキーを含みます。
messages: OpenAI messages形式の対話データ
instruction_1: 1ターン目の指示プロンプト
output_1: 1ターン目のアシスタント応答
instruction_2: 2ターン目の指示プロンプト
output_2: 2ターン目のアシスタント応答
1ターン目の指示プロンプトはdeepseek-ai/DeepSeek-V3-0324で合成されています。system promptや2ターン目の指示プロンプトは事前に用意した複数種類からランダムに選択されたものが設定されています。
ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-Creative-Writing-39.6k.ja_asr.jsut_basic5000whisper_transcriptions.mlsJapanese-Novels-23M
Japanese-Novels-23M
This dataset contains Japanese web novels that I collected personally.
Machine-Learning Use OnlyAccess is restricted to bona fide machine-learning–related purposes.To request access, please provide a detailed explanation of the specific tasks or applications for which you intend to use the dataset.
Total records: 23,212,809
Total characters: 80,846,120,027
Total tokens (Llama 4 tokenizer): 55,406,468,406 (55.4 B)
sentence_transformer_japanese
日本語のデータセットを SentenceTransformes で学習しやすいカラム名と構造に変換したもの。
主に (anchor, positive), (anchor, positive, negative), (anchor, positive, negative_1, ..., negative_n) といった構造になっているため、とりわけ対照学習で使いやすくなっています。
以下のデータセットから作成
https://huggingface.co/datasets/hpprc/emb
https://huggingface.co/datasets/hotchpotch/hpprc_emb-scores のリランカースコアを用いて、positive(>=0.7) / negative(<=0.3) のフィルタリングを行った
https://huggingface.co/datasets/hpprc/llmjp-kaken
https://huggingface.co/datasets/hpprc/msmarco-ja… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/sentence_transformer_japanese.laion2B-japanese-subsetawesome-japanese-corpus
Awesome Japanese Corpus
4つの日本語データソースを text、from、from_license の3列に正規化した
Parquet データセットです。
FineWeb以外の3ソースは、生成処理を停止した時点までに取得済みの公開データを
採用しています。FineWebは最大3シャードを並列先読みする1時間限定の処理で
サンプリングしています。空文字は除外しています。Infini-News は取得済みの
年度について language_iso639_3 == "jpn" の行を採用しています。
Sources
hotchpotch/fineweb-2-edu-japanese (odc-by)
ruggsea/infini-news-corpus の language_iso639_3 == "jpn" (cc-by-4.0)
turing-motors/MOMIJI (cc-by-4.0)
AhmedSSabir/Japanese-wiki-dump-sentence-dataset… See the full description on the dataset page: https://huggingface.co/datasets/nakasyou/awesome-japanese-corpus.whisper_transcriptions.reazonspeech.large.wer_10.0Japanese_Bandori_Band_Story
Japanese Bandori Band Story
Japanese Band Story text retrieved from the Bestdori scenario assets.
This snapshot contains 26 story entries, 493 chapters,
and 30679 rows (28800 dialogue rows).
Created at 2026-09-15T02:11:27.707570+00:00.
Files
data/train-*.parquet: Hub dataset shards generated by Dataset.push_to_hub.
data/band_stories.jsonl: local combined dataset, also included in the downloadable ZIP.
stories/story_XXXX/: complete per-story TXT, CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/KomeijiForce/Japanese_Bandori_Band_Story.Japanese-Eroge-Voice
Japanese-Eroge-Voice
Description
This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model.
Preprocessing Steps
The raw audio data has undergone the following preprocessing steps:
Loudness Normalization:
Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice.Japanese-Eroge-Voice-V2
Japanese-Eroge-Voice-V2
Description
This is the successor to the Japanese-Eroge-Voice dataset. It consists of a significantly larger collection of audio-transcription pairs extracted from Japanese eroge (adult games).
Note on Versioning: There is no overlap between this dataset (V2) and the previous version. All audio clips and transcriptions in V2 are distinct from those in the original version, providing entirely new data for research.
This version (V2) expands the… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice-V2.japanese-anime-speech-v2
Japanese Anime Speech Dataset V2
日本語はこちら
japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models.
The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels.
This dataset is not an updated version of japanese-anime-speech-v1.
For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset.
The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/WatsonNT/japanese-anime-speech-v2.japanese-conversion
Awesome Japanese IME Training Data
Awesome Japanese Corpus の本文を直接 KyTea で解析し、文脈付きかな漢字変換の
ランキング学習例を作成したデータセットです。中間の読み付きデータセットは
作りません。任意の検証モードでは、抽出範囲についてMeCabの読みとも一致した
例だけを採用できます。
context: 変換対象より前の本文
input: 変換対象のひらがな読み
correct: 元コーパスにある正解表記
incorrect: predict.py で全体または一部分を再変換した誤候補の配列
n_words: 抽出した連続形態素数
source_text と target_start / target_end により、元文章中の抽出位置を
復元できます。元データの利用条件は from と from_license を参照して
ください。
whisper_transcriptions.reazonspeech.mediumJapanese-PD
🇯🇵 Japanese Public Domain 🇯🇵
Japanese-Public Domain or Japanese-PD is a large collection aiming to aggregate all Japanese monographies and periodicals in the public domain.
Dataset summary
The collection contains 1,410 titles making up 21,072,188 words recovered from the Internet Archive. Each parquet file has the full text of 2,000 books selected at random.
Curation method
The composition of the dataset adheres to the criteria for public domain works in… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Japanese-PD.japanese-context-relevanceこのデータセットは抽出されたハードネガティブ、質問とテキストをさらに細かく区切ったspanとの関連度スコア、さらにリランカーのスコアが含まれており、OpenProvence などのモデル学習に利用できます。
サブセットごとに元データが異なるため、ライセンスはそれぞれの提供元に従ってください。
利用可能なサブセット一覧
重複除去の有無ごとにサブセット構成をまとめました。freq2 系は MD5 ベースの頻度フィルタでデータセット全体に同一テキストが 3 回以上出現しないよう調整しており、軽量でバランスの良い学習データが欲しい場合はこちらを推奨します。同じテキストが繰り返し登場すると context_spans_relevance が過学習しやすくなるため、剪定モデルの訓練では重複を抑えることを推奨します。
サブセット
行数 (train / val / test)
テキスト重複率
推奨
msmarco-ja
492,729 / 5,000 / 5,000
約 38.9%
msmarco-ja-freq2
260,436 / 1,000… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/japanese-context-relevance.whisper_transcriptions.reazonspeech.largejapanese-anime-speech-v2-split-150k
japanese-anime-speech-v2-split-150k
joujiboi/japanese-anime-speech-v2 的 150,000 筆隨機子集,已切好 train / test。
split
rows
train
135,000
test
15,000
total
150,000
Columns
audio — 16 kHz mp3,與原始資料完全相同(未重新編碼)
sentence — 轉錄文字(原始欄位名為 transcription)
How it was built
來源的 sfw(271,788 筆)與 nsfw(20,849 筆)兩個 split 都有使用,並依原始比例分配名額
(sfw 139,313 / nsfw 10,687)。
在每個 split 的全部列上做無放回均勻抽樣,因此每筆資料被選中的機率相同。
抽出後整體打亂,前 15,000 筆為 test,其餘為 train,train / test… See the full description on the dataset page: https://huggingface.co/datasets/hhim8826/japanese-anime-speech-v2-split-150k.llm-japanese-dataset
llm-japanese-dataset
LLM構築用の日本語インストラクション(チャット)データセット
主に,英語で構築されたLLMモデルなどに対して,チャット(Instruction)応答タスクに関してLoRAなどでチューニングするために使用できます.
※様々な公開言語資源を利用させていただきました.関係各位にはこの場を借りて御礼申し上げます.
updates
2023/5/15にAlpaca datasetがNCにライセンス変更されたことに対応し,安心してご利用いただけるように,データセットから当該データセットをドロップしました.
v1.0.1にて,ドロップ後のデータセットをご利用いただけます.
2024/1/4にWikipedia summaryに空白文字のみで構成される出力を削除することに対応し,Wikipediaのバージョンアップデート(20240101)をしました(v1.0.2).
2024/1/18にAsian Language Treebank (ALT)データセットの欠損した出力を削除しました(v1.0.3).… See the full description on the dataset page: https://huggingface.co/datasets/izumi-lab/llm-japanese-dataset.
