CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01japanese-asr /whisper_transcriptions.reazon_speech_allaudio10M<n<100M16 likes47k downloads2y agoHugging Face02japanese-asr /whisper_transcriptions.reazonspeech.allaudio10M<n<100M4 likes7.3k downloads2y agoHugging Face03japanese-asr /ja_asr.reazon_speech_allaudio10M<n<100M7 likes5.3k downloads2y agoHugging Face04joujiboi /japanese-anime-speech-v2 Japanese Anime Speech Dataset V2 日本語はこちら japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models. The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels. This dataset is not an updated version of japanese-anime-speech-v1. For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset. The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2.audioautomatic-speech-recognition100K<n<1M153 likes4.9k downloads11mo agoHugging Face05japanese-asr /whisper_transcriptions.reazonspeech.all.wer_10.0audio1M<n<10M3 likes3.5k downloads2y agoHugging Face06japanese-asr /en_asr.mlsaudio10M<n<100M3 likes3.3k downloads2y agoHugging Face07hotchpotch /fineweb-2-edu-japanese 🍷 FineWeb2 Edu Japanese: High-Quality Educational Japanese Dataset This dataset consists of 120 million texts (approximately 89.3B tokens) filtered from the 376 million Japanese texts in FineWeb2 that were deemed educational. The following subsets are also provided: default: Approximately 120M texts (120 million texts) totaling around 89.3B tokens sample_10BT: A random sample of about 10B tokens from the default dataset small_tokens: Data composed solely of texts with 512 tokens… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-2-edu-japanese.tabular100M<n<1B34 likes3.3k downloads1y agoHugging Face08daisuke9999 /Japanese_NicoNico_Douga_Movie_Meta_Data_2016tabular10M<n<100M0 likes3k downloads2y agoHugging Face09joujiboi /japanese-anime-speech Japanese Anime Speech Dataset 日本語はこちら japanese-anime-speech is an audio-text dataset designed for the training of automatic speech recognition models. The dataset is comprised of thousands of audio clips and their corresponding transcriptions from different visual novels. The goal of this dataset is to increase the accuracy of automatic speech recognition models, such as OpenAI's Whisper, in accurately transcribing dialogue from anime and other similar Japanese media. This genre is… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/japanese-anime-speech.audioautomatic-speech-recognition10K<n<100K161 likes1.5k downloads2y agoHugging Face10MIL-UT /Japanese-Medical-VQA-12m Japanese Medical VQA 12M Japanese Medical VQA 12M is a large-scale Japanese medical multimodal dataset built from Open-PMC-18M and released in Parquet and Webdataset format. This dataset contains outputs from multiple data-construction stages, including: source captions Japanese translations of source captions enriched captions Japanese translations of enriched captions question-answering Current Repository Format This repository currently stores the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/MIL-UT/Japanese-Medical-VQA-12m.imageimage-to-text10M<n<100M7 likes1.3k downloads7mo agoHugging Face11oshizo /japanese-text-image-retrieval-trainshunk031/JDocQAのtrain splitに含まれるPDFデータを画像化し、NDLOCRでOCRしたテキストとペアにしたデータセットです。OCRは長い辺を1200pxにリサイズした画像に対して実施しました。OCR結果には、読み取りに失敗した際の文字列「〓」が含まれます。本データセットに含めている画像は、長い辺を896px、700px、588pxのいずれかにリサイズしています。どのサイズとするかは主にページに含まれる文字数で決めました。 query列は、OCR結果の文字列に対しQwen/Qwen2.5-14B-Instructで生成したものです。3つの質問を生成させ、ランダムに1つを選んだものをデータセットに含めました。質問を生成する際は以下のプロンプトを使用しました。 あなたは、質問から画像をretrieveするためのモデルをトレーニングするための(質問, 画像)ペアのデータセットを作成するプロジェクトのメンバーである。 プロジェクトは以下のように進める。 step1. ドキュメントPDFを1ページ1枚の画像ファイルに変換する step2.… See the full description on the dataset page: https://huggingface.co/datasets/oshizo/japanese-text-image-retrieval-train.image100K<n<1M0 likes1.1k downloads2y agoHugging Face12Aratako /Japanese-Creative-Writing-39.6k Japanese-Creative-Writing-39.6k 概要 deepseek-ai/DeepSeek-V3-0324を用いて作成した、約39600件の日本語の小説執筆タスクデータセットです。 全てのデータは2ターンのデータとなっています。また、データセット中の一部データはNSFW表現を含みます。 データの詳細 各データは以下のキーを含みます。 messages: OpenAI messages形式の対話データ instruction_1: 1ターン目の指示プロンプト output_1: 1ターン目のアシスタント応答 instruction_2: 2ターン目の指示プロンプト output_2: 2ターン目のアシスタント応答 1ターン目の指示プロンプトはdeepseek-ai/DeepSeek-V3-0324で合成されています。system promptや2ターン目の指示プロンプトは事前に用意した複数種類からランダムに選択されたものが設定されています。 ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-Creative-Writing-39.6k.texttext-generation10K<n<100K8 likes1k downloads1y agoHugging Face13japanese-asr /ja_asr.jsut_basic5000audio1K<n<10K11 likes920 downloads2y agoHugging Face14japanese-asr /whisper_transcriptions.mlsaudio10M<n<100M1 likes918 downloads2y agoHugging Face15OmniAICreator /Japanese-Novels-23Mgated Japanese-Novels-23M This dataset contains Japanese web novels that I collected personally. Machine-Learning Use OnlyAccess is restricted to bona fide machine-learning–related purposes.To request access, please provide a detailed explanation of the specific tasks or applications for which you intend to use the dataset. Total records: 23,212,809 Total characters: 80,846,120,027 Total tokens (Llama 4 tokenizer): 55,406,468,406 (55.4 B) tabulartext-generation10M<n<100M29 likes911 downloads1y agoHugging Face16hotchpotch /sentence_transformer_japanese 日本語のデータセットを SentenceTransformes で学習しやすいカラム名と構造に変換したもの。 主に (anchor, positive), (anchor, positive, negative), (anchor, positive, negative_1, ..., negative_n) といった構造になっているため、とりわけ対照学習で使いやすくなっています。 以下のデータセットから作成 https://huggingface.co/datasets/hpprc/emb https://huggingface.co/datasets/hotchpotch/hpprc_emb-scores のリランカースコアを用いて、positive(>=0.7) / negative(<=0.3) のフィルタリングを行った https://huggingface.co/datasets/hpprc/llmjp-kaken https://huggingface.co/datasets/hpprc/msmarco-ja… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/sentence_transformer_japanese.text10M<n<100M7 likes904 downloads2y agoHugging Face17lyakaap /laion2B-japanese-subsetimage100M<n<1B4 likes848 downloads4y agoHugging Face18nakasyou /awesome-japanese-corpus Awesome Japanese Corpus 4つの日本語データソースを text、from、from_license の3列に正規化した Parquet データセットです。 FineWeb以外の3ソースは、生成処理を停止した時点までに取得済みの公開データを 採用しています。FineWebは最大3シャードを並列先読みする1時間限定の処理で サンプリングしています。空文字は除外しています。Infini-News は取得済みの 年度について language_iso639_3 == "jpn" の行を採用しています。 Sources hotchpotch/fineweb-2-edu-japanese (odc-by) ruggsea/infini-news-corpus の language_iso639_3 == "jpn" (cc-by-4.0) turing-motors/MOMIJI (cc-by-4.0) AhmedSSabir/Japanese-wiki-dump-sentence-dataset… See the full description on the dataset page: https://huggingface.co/datasets/nakasyou/awesome-japanese-corpus.texttext-generation100M<n<1B1 likes844 downloads2mo agoHugging Face19japanese-asr /whisper_transcriptions.reazonspeech.large.wer_10.0audio1M<n<10M0 likes665 downloads3y agoHugging Face20KomeijiForce /Japanese_Bandori_Band_Story Japanese Bandori Band Story Japanese Band Story text retrieved from the Bestdori scenario assets. This snapshot contains 26 story entries, 493 chapters, and 30679 rows (28800 dialogue rows). Created at 2026-09-15T02:11:27.707570+00:00. Files data/train-*.parquet: Hub dataset shards generated by Dataset.push_to_hub. data/band_stories.jsonl: local combined dataset, also included in the downloadable ZIP. stories/story_XXXX/: complete per-story TXT, CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/KomeijiForce/Japanese_Bandori_Band_Story.tabulartext-generation10K<n<100K0 likes601 downloads11d agoHugging Face21NandemoGHS /Japanese-Eroge-Voice Japanese-Eroge-Voice Description This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model. Preprocessing Steps The raw audio data has undergone the following preprocessing steps: Loudness Normalization: Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice.audiotext-to-speech100K<n<1M37 likes597 downloads1y agoHugging Face22NandemoGHS /Japanese-Eroge-Voice-V2 Japanese-Eroge-Voice-V2 Description This is the successor to the Japanese-Eroge-Voice dataset. It consists of a significantly larger collection of audio-transcription pairs extracted from Japanese eroge (adult games). Note on Versioning: There is no overlap between this dataset (V2) and the previous version. All audio clips and transcriptions in V2 are distinct from those in the original version, providing entirely new data for research. This version (V2) expands the… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice-V2.audiotext-to-speech1M<n<10M53 likes585 downloads8mo agoHugging Face23WatsonNT /japanese-anime-speech-v2 Japanese Anime Speech Dataset V2 日本語はこちら japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models. The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels. This dataset is not an updated version of japanese-anime-speech-v1. For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset. The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/WatsonNT/japanese-anime-speech-v2.audioautomatic-speech-recognition100K<n<1M1 likes573 downloads29d agoHugging Face24nakasyou /japanese-conversion Awesome Japanese IME Training Data Awesome Japanese Corpus の本文を直接 KyTea で解析し、文脈付きかな漢字変換の ランキング学習例を作成したデータセットです。中間の読み付きデータセットは 作りません。任意の検証モードでは、抽出範囲についてMeCabの読みとも一致した 例だけを採用できます。 context: 変換対象より前の本文 input: 変換対象のひらがな読み correct: 元コーパスにある正解表記 incorrect: predict.py で全体または一部分を再変換した誤候補の配列 n_words: 抽出した連続形態素数 source_text と target_start / target_end により、元文章中の抽出位置を 復元できます。元データの利用条件は from と from_license を参照して ください。 tabulartext-generation10M<n<100M1 likes550 downloads2mo agoHugging Face25japanese-asr /whisper_transcriptions.reazonspeech.mediumaudio100K<n<1M0 likes512 downloads3y agoHugging Face26PleIAs /Japanese-PD 🇯🇵 Japanese Public Domain 🇯🇵 Japanese-Public Domain or Japanese-PD is a large collection aiming to aggregate all Japanese monographies and periodicals in the public domain. Dataset summary The collection contains 1,410 titles making up 21,072,188 words recovered from the Internet Archive. Each parquet file has the full text of 2,000 books selected at random. Curation method The composition of the dataset adheres to the criteria for public domain works in… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Japanese-PD.tabular1M<n<10M1 likes492 downloads7mo agoHugging Face27hotchpotch /japanese-context-relevanceこのデータセットは抽出されたハードネガティブ、質問とテキストをさらに細かく区切ったspanとの関連度スコア、さらにリランカーのスコアが含まれており、OpenProvence などのモデル学習に利用できます。 サブセットごとに元データが異なるため、ライセンスはそれぞれの提供元に従ってください。 利用可能なサブセット一覧 重複除去の有無ごとにサブセット構成をまとめました。freq2 系は MD5 ベースの頻度フィルタでデータセット全体に同一テキストが 3 回以上出現しないよう調整しており、軽量でバランスの良い学習データが欲しい場合はこちらを推奨します。同じテキストが繰り返し登場すると context_spans_relevance が過学習しやすくなるため、剪定モデルの訓練では重複を抑えることを推奨します。 サブセット 行数 (train / val / test) テキスト重複率 推奨 msmarco-ja 492,729 / 5,000 / 5,000 約 38.9% msmarco-ja-freq2 260,436 / 1,000… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/japanese-context-relevance.text1M<n<10M0 likes484 downloads11mo agoHugging Face28japanese-asr /whisper_transcriptions.reazonspeech.largeaudio1M<n<10M0 likes446 downloads3y agoHugging Face29hhim8826 /japanese-anime-speech-v2-split-150k japanese-anime-speech-v2-split-150k joujiboi/japanese-anime-speech-v2 的 150,000 筆隨機子集,已切好 train / test。 split rows train 135,000 test 15,000 total 150,000 Columns audio — 16 kHz mp3,與原始資料完全相同(未重新編碼) sentence — 轉錄文字(原始欄位名為 transcription) How it was built 來源的 sfw(271,788 筆)與 nsfw(20,849 筆)兩個 split 都有使用,並依原始比例分配名額 (sfw 139,313 / nsfw 10,687)。 在每個 split 的全部列上做無放回均勻抽樣,因此每筆資料被選中的機率相同。 抽出後整體打亂,前 15,000 筆為 test,其餘為 train,train / test… See the full description on the dataset page: https://huggingface.co/datasets/hhim8826/japanese-anime-speech-v2-split-150k.audioautomatic-speech-recognition100K<n<1M0 likes434 downloads2mo agoHugging Face30izumi-lab /llm-japanese-dataset llm-japanese-dataset LLM構築用の日本語インストラクション(チャット)データセット 主に,英語で構築されたLLMモデルなどに対して,チャット(Instruction)応答タスクに関してLoRAなどでチューニングするために使用できます. ※様々な公開言語資源を利用させていただきました.関係各位にはこの場を借りて御礼申し上げます. updates 2023/5/15にAlpaca datasetがNCにライセンス変更されたことに対応し,安心してご利用いただけるように,データセットから当該データセットをドロップしました. v1.0.1にて,ドロップ後のデータセットをご利用いただけます. 2024/1/4にWikipedia summaryに空白文字のみで構成される出力を削除することに対応し,Wikipediaのバージョンアップデート(20240101)をしました(v1.0.2). 2024/1/18にAsian Language Treebank (ALT)データセットの欠損した出力を削除しました(v1.0.3).… See the full description on the dataset page: https://huggingface.co/datasets/izumi-lab/llm-japanese-dataset.text1M<n<10M144 likes415 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.