datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fungi_diagnostic_chars_comparison_japanese
fungi_diagnostic_chars_comparison_japanese大菌輪「識別形質まとめ」データセット最終更新日 / Last updated: 2026/8/29(up to R3-14214)
Languages
Japanese
This dataset is available in Japanese only.
概要 / Overview
Atsushi Nakajima(中島淳志)が個人で運営しているWebサイト大菌輪では、数千件以上の菌類分類学論文を「論文3行まとめ」という形で要約および索引付け(インデキシング)した情報を提供しています。その一環として、ある菌と別の菌の「共通する」あるいは「異なる」識別形質 (diagnostic characters) に関する記述を人手で抽出しています。
Daikinrin, a personal website run by Atsushi Nakajima, provides summaries and… See the full description on the dataset page: https://huggingface.co/datasets/Atsushi/fungi_diagnostic_chars_comparison_japanese.japanese-image-classification-evaluation-dataset
recruit-jp/japanese-image-classification-evaluation-dataset
Overview
Developed by: Recruit Co., Ltd.
Dataset type: Image Classification
Language(s): Japanese
LICENSE: CC-BY-4.0
More details are described in our tech blog post.
日本語CLIP学習済みモデルとその評価用データセットの公開
Dataset Details
This dataset is comprised of four image classification tasks related to concepts and things unique to Japan. Specifically, is consists of the following tasks.
jafood101: Image… See the full description on the dataset page: https://huggingface.co/datasets/recruit-jp/japanese-image-classification-evaluation-dataset.japanese-speech-recognition-dataset
Japanese Speech Dataset for recognition task
Dataset comprises 10+ hours of telephone dialogues in Japanese, collected from 10 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/japanese-speech-recognition-dataset.JapaneseWordsDictionaryJParaCrawl-Filtered-English-Japanese-Parallel-Corpus
Introduction
This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus.
The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet.
Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.Japanese-Speech-Dataset
🎧 Japanese Speech Dataset
The Japanese Speech Dataset is a production-ready speech audio dataset designed to provide high-quality, structured audio data for AI and machine learning applications. It includes 132 hours of audio data distributed across 733 files, available in MP3 and WAV formats, with a total size of 272 MB. This well-balanced audio dataset delivers diverse voice data, with 54% female and 46% male speakers, and an age distribution spanning from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Japanese-Speech-Dataset.Japanese_sentimentjapanese-character-difficulty
Japanese Character Difficulty Dataset
A comprehensive dataset of 3,003 Japanese kanji characters with their educational difficulty grades, sourced from official Japanese educational standards and kanjiapi.dev.
Dataset Overview
Total Characters: 3,003 kanji
Source: Japanese Ministry of Education (MEXT) Joyo Kanji list + kanjiapi.dev
Coverage: Elementary grades 1-6, plus secondary education and advanced characters
Format: Character-grade pairs for easy lookup and analysis… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-character-difficulty.three_line_summarization_for_japanese_news_articlesライブドアニュースコーパスの3行要約データセットです。
Llama v2向けのプロンプトを追加して成形してあります。
学習に利用する際は、 [R_START] [R_END] をspecial tokenとして追加することを推奨します。
Number of rows: 3,907
Datasetは以下のリポジトリを利用してscrapeしました。
git@github.com:KodairaTomonori/ThreeLineSummaryDataset.git
chatgpt-prompts-Japanesejapanese-names-datasetjapanese-grammar-correction
Japanese Grammar Correction Dataset
Description
This dataset is designed to train language models to identify and correct a wide range of grammatical errors and stylistic issues in Japanese text.
The data consists of pairs of incorrect and correct sentences, along with metadata that classifies the type of error and provides additional context.
The dataset was created by both manual curation from discussions in Japanese learning communities and synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/huytd189/japanese-grammar-correction.japanesejapanese_recipe本データセットは、産業技術大学院大学(AIIT)在学中に完成したものです。
英語で書いてあるレシピデータセットを日本語に翻訳する。
元データセット:Shengtao/recipe
データ数:3757
这个数据集是将原来英语的菜谱数据集翻译为日语。
原数据集:https://huggingface.co/datasets/Shengtao/recipe
数据量:3757条
OmniL2L-Myanmar-Japanese
🇲🇲 🔄 🇯🇵 OmniL2L-Myanmar-Japanese
💡 Note: This dataset is a dedicated language-pair component of the main multi-language corpus. To access the complete multi-lingual matrix combining all languages simultaneously, please visit the main repository: kalixlouiis/OmniL2L.
OmniL2L-Myanmar-Japanese is a trustworthy, human-verified parallel translation dataset pairing Burmese (Myanmar) with Japanese. This dataset is custom-tailored for low-resource machine translation and… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/OmniL2L-Myanmar-Japanese.Japanese-English_translation_of_contents_HScodes日本郵便が提供する「国際郵便 内容品の日英・中英訳、HSコード類」(2024/05/09)のデータに基づいています。
詳しくはサイトをご覧ください
https://www.post.japanpost.jp/int/use/publication/contentslist/index.php?id=0&ie=utf8&lang=_ja&q=
japanese_relation_triple_datasetJapanese_NER_Data_Hub
概要
大規模言語モデル(LLM)用の固有表現認識データセット(J-NER)のリポジトリです。
J-NERは拡張固有表現階層(*)の内、応用の観点から重要な名前表現から構成される総計157種類の固有表現を含んだデータセットです。
J-NERで取り扱う固有表現はLLMの学習データに含まれていることが要求されるため、J-NERに含まれる固有表現はWikipediaにページが存在する単語のみとしています。
各固有表現に関して、その固有表現を含んだデータ(文)である正例5例、含まれていないデータ(文)である負例5例がデータセットに存在します。
よってデータセットの総計は157×5×2=1,570です。
*:2024年5月28日現在、拡張固有表現階層の以下サイトにアクセスできません。
https://ene-project.info/ene9/
使い方
データロード
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/sergicalsix/Japanese_NER_Data_Hub.myanmar-japanesejapanese-speech-recognition-dataset
Japanese Telephone Dialogues Dataset - 10 Hours
Dataset comprises 10 hours of high-quality telephone audio recordings in Japanese, featuring 20+ native speakers and achieving a 95% sentence accuracy rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/japanese-speech-recognition-dataset.multils-japanese
MultiLS-Japanese
MultiLS-Japanese is a lexical complexity prediction (LCP) and lexical simplification (LS) dataset for Japanese. The MultiLS-Japanese dataset was created by Adam Nohejl, Akio Haykawa, and Yusuke Ide.
A journal paper about the dataset.
More information and additional files on the MultiLS-Japanese Github repo.
Multilingual LS and LCP Data
Related datasets for 9 more languages: MLSP2024 dataset on Hugging Face Hub.
mental_health_japanese_data
Mental Health Dataset (Japanese)
Dataset Summary
This dataset was independently created during my study at the Advanced Institute of Industrial Technology (AIIT) for research and educational purposes.It is designed to support natural language processing (NLP) tasks related to mental health dialogues and response generation.
Dataset Structure
The dataset consists of two main columns:
input: user-side utterances or questions
output: system-side supportive… See the full description on the dataset page: https://huggingface.co/datasets/xiashuaxia/mental_health_japanese_data.naga-eng-fullJapanese-Complex-GraphQA
Japanese-Complex-GraphQA
Japanese-Complex-GraphQA は,日本語の図表(チャート・グラフ)を対象とした図表質問応答(Graph Question Answering)データセットです。高度な読解・比較・計算を必要とする難易度の高い質問を含むことを特徴としています。
本データセットは,図表を含む公開統計資料を基に,質問文おび正答を新たに作成し,日本語環境におけるマルチモーダル QA および推論能力の評価を目的として構築されました。
データセット構成
本データセットは以下のファイルから構成されます。
ファイル一覧
jc_graphqa.csv質問応答データ(id, 質問文, 正答, 画像ファイルid)
jc_graphqa_tags.jsonl各質問に付与された推論・図表属性タグ
CSV ファイル形式(data.csv)
カラム名
内容
id
データ識別子(数値)
question
日本語の質問文
answer
正答… See the full description on the dataset page: https://huggingface.co/datasets/ab528/Japanese-Complex-GraphQA.JapanesePhisingjapanese_fake_newsjapanese_summaryJapanese-Vocab-6knaga-eng
