CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hotchpotch /fineweb-2-edu-japanese 🍷 FineWeb2 Edu Japanese: High-Quality Educational Japanese Dataset This dataset consists of 120 million texts (approximately 89.3B tokens) filtered from the 376 million Japanese texts in FineWeb2 that were deemed educational. The following subsets are also provided: default: Approximately 120M texts (120 million texts) totaling around 89.3B tokens sample_10BT: A random sample of about 10B tokens from the default dataset small_tokens: Data composed solely of texts with 512 tokens… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-2-edu-japanese.tabular100M<n<1B34 likes3.3k downloads1y agoHugging Face02daisuke9999 /Japanese_NicoNico_Douga_Movie_Meta_Data_2016tabular10M<n<100M0 likes3k downloads2y agoHugging Face03OmniAICreator /Japanese-Novels-23Mgated Japanese-Novels-23M This dataset contains Japanese web novels that I collected personally. Machine-Learning Use OnlyAccess is restricted to bona fide machine-learning–related purposes.To request access, please provide a detailed explanation of the specific tasks or applications for which you intend to use the dataset. Total records: 23,212,809 Total characters: 80,846,120,027 Total tokens (Llama 4 tokenizer): 55,406,468,406 (55.4 B) tabulartext-generation10M<n<100M29 likes914 downloads1y agoHugging Face04lyakaap /laion2B-japanese-subsetimage100M<n<1B4 likes834 downloads4y agoHugging Face05KomeijiForce /Japanese_Bandori_Band_Story Japanese Bandori Band Story Japanese Band Story text retrieved from the Bestdori scenario assets. This snapshot contains 26 story entries, 493 chapters, and 30679 rows (28800 dialogue rows). Created at 2026-09-15T02:11:27.707570+00:00. Files data/train-*.parquet: Hub dataset shards generated by Dataset.push_to_hub. data/band_stories.jsonl: local combined dataset, also included in the downloadable ZIP. stories/story_XXXX/: complete per-story TXT, CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/KomeijiForce/Japanese_Bandori_Band_Story.tabulartext-generation10K<n<100K0 likes604 downloads12d agoHugging Face06nakasyou /japanese-conversion Awesome Japanese IME Training Data Awesome Japanese Corpus の本文を直接 KyTea で解析し、文脈付きかな漢字変換の ランキング学習例を作成したデータセットです。中間の読み付きデータセットは 作りません。任意の検証モードでは、抽出範囲についてMeCabの読みとも一致した 例だけを採用できます。 context: 変換対象より前の本文 input: 変換対象のひらがな読み correct: 元コーパスにある正解表記 incorrect: predict.py で全体または一部分を再変換した誤候補の配列 n_words: 抽出した連続形態素数 source_text と target_start / target_end により、元文章中の抽出位置を 復元できます。元データの利用条件は from と from_license を参照して ください。 tabulartext-generation10M<n<100M1 likes588 downloads2mo agoHugging Face07llm-jp /relaion2B-en-research-safe-japanese-translation relaion2B-en-research-safe-japanese-translation This dataset is the Japanese translation of the English subset of ReLAION-5B (laion/relaion2B-en-research-safe), translated by gemma-2-9b-it. We used text2dataset for translating with open-weight LLMs. By leveraging the fast LLM inference library vLLM, this tool enables the rapid translation of large English datasets into Japanese. Prompt The following is the prompt used for translation with Gemma. You are an… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/relaion2B-en-research-safe-japanese-translation.image1B<n<10B4 likes519 downloads1y agoHugging Face08PleIAs /Japanese-PD 🇯🇵 Japanese Public Domain 🇯🇵 Japanese-Public Domain or Japanese-PD is a large collection aiming to aggregate all Japanese monographies and periodicals in the public domain. Dataset summary The collection contains 1,410 titles making up 21,072,188 words recovered from the Internet Archive. Each parquet file has the full text of 2,000 books selected at random. Curation method The composition of the dataset adheres to the criteria for public domain works in… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Japanese-PD.tabular1M<n<10M1 likes478 downloads7mo agoHugging Face09ShrimpLM /fineweb-2-edu-japanese-top30mhotchpotch/fineweb-2-edu-japanese こちらのデータセットのscoreが3.2以上のデータを抽出したものです。 良質なデータを公開してくださったhotchpotch様に感謝申し上げます。 License This dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0, as is the original FineWeb2 dataset. Additionally, its use is subject to the CommonCrawl Terms of Use. tabular10M<n<100M0 likes254 downloads2mo agoHugging Face10Atsushi /fungi_diagnostic_chars_comparison_japanese fungi_diagnostic_chars_comparison_japanese大菌輪「識別形質まとめ」データセット最終更新日 / Last updated: 2026/8/29(up to R3-14214) Languages Japanese This dataset is available in Japanese only. 概要 / Overview Atsushi Nakajima(中島淳志)が個人で運営しているWebサイト大菌輪では、数千件以上の菌類分類学論文を「論文3行まとめ」という形で要約および索引付け(インデキシング)した情報を提供しています。その一環として、ある菌と別の菌の「共通する」あるいは「異なる」識別形質 (diagnostic characters) に関する記述を人手で抽出しています。 Daikinrin, a personal website run by Atsushi Nakajima, provides summaries and… See the full description on the dataset page: https://huggingface.co/datasets/Atsushi/fungi_diagnostic_chars_comparison_japanese.tabulartext-classification100K<n<1M0 likes192 downloads28d agoHugging Face11answerdotai /MMARCO-japanese-32-scored-triplets@misc{clavié2024jacolbertv25optimisingmultivectorretrievers, title={JaColBERTv2.5: Optimising Multi-Vector Retrievers to Create State-of-the-Art Japanese Retrievers with Constrained Resources}, author={Benjamin Clavié}, year={2024}, eprint={2407.20750}, archivePrefix={arXiv}, primaryClass={cs.IR}, url={https://arxiv.org/abs/2407.20750}, } tabular1M<n<10M6 likes184 downloads2y agoHugging Face12bowang0911 /japanese-conala Attribution MTEB-format derivative of haih2/japanese-conala (Japanese CoNaLa; intent->code). Query = Japanese intent; corpus = Python snippet. Paired test+train combined. tabulartext-retrieval1K<n<10K0 likes148 downloads3mo agoHugging Face13mini97 /filtered_japanese-wikipediatabular1M<n<10M2 likes144 downloads1y agoHugging Face14inarikami /wikipedia-japanese Japanese Wikipedia Dataset This dataset is a comprehensive pull of all Japanese wikipedia article data as of 20220808. Note: Right now its uploaded as a single cleaned gzip file (for faster usage), I'll update this in the future to include a huggingface datasets compatible class and better support for japanese than the existing wikipedia repo. Example use case: gunzip jawwiki20200808.json.gz import pandas as pd from datasets import load_dataset df =… See the full description on the dataset page: https://huggingface.co/datasets/inarikami/wikipedia-japanese.tabular1M<n<10M5 likes119 downloads4y agoHugging Face15oshizo /japanese-wikipedia-paragraphsA slightly modified version of the parsing and chunking method for singletongue/wikipedia-utils. Pre-processing was performed using oshizo/wikipedia-utils, which is a fork of the original repository, singletongue/wikipedia-utils. The Wikipedia data was crawled between 2023/12/5 and 2023/12/8. tabular10M<n<100M4 likes115 downloads3y agoHugging Face16llm-book /aio-passages-bpr-bert-base-japanese-v3 Dataset Card for llm-book/aio-passages-bert-base-japanese-v3-bpr 書籍『大規模言語モデル入門』で使用する、「AI王」コンペティションのパッセージデータセットに BPR によるパッセージの埋め込みを適用したデータセットです。 llm-book/aio-passages のデータセットに対して、llm-book/bert-base-japanese-v3-bpr-passage-encoder によるパッセージのバイナリベクトルが embeddings フィールドに追加されています。 Licence 本データセットで利用している Wikipedia のコンテンツは、クリエイティブ・コモンズ表示・継承ライセンス 3.0 (CC BY-SA 3.0) および GNU 自由文書ライセンス (GFDL) の下に配布されているものです。 tabular1M<n<10M1 likes110 downloads3y agoHugging Face17hotchpotch /fineweb-2-edu-japanese-noise-detect-rawfineweb-2-edu-japanese の small_tokens の text カラムをユニコード正規化(NFKC)したものを fineweb-2-japanese-text-cleaner を使ってノイズ箇所を推論したRAWデータセットです。 このデータセットで、ノイズ文字列を削除したのものを、fineweb-2-edu-japaneseのsmall_tokens_cleanedサブセットとして公開しています。 推論時のパラメータは閾値0.7以上かつノイズの文字列長4文字以上のものを、noise_spans カラムに付与しています。noise_spans は start_pos, end_pos のペアとなってます。 ライセンス ODC-By tabular10M<n<100M0 likes105 downloads2y agoHugging Face18lynx1231 /historical-japanese-yen-cross-forwards-sample Historical Japanese Yen Cross FX Forwards Sample A free evaluation sample of historical FX-forward data for selected non-USD Japanese yen crosses across multiple tenors. Full historical FX-forward catalog, broader pair and tenor coverage, downloadable datasets, and pricing:https://futuresforexandsomeindexes.com/ This repository is a free evaluation sample intended for schema inspection, data-quality evaluation, integration testing, and quantitative research prototyping. It is… See the full description on the dataset page: https://huggingface.co/datasets/lynx1231/historical-japanese-yen-cross-forwards-sample.tabular1M<n<10M0 likes100 downloads8d agoHugging Face19Kotomiya07 /premodern-japanese-books-lm-corpus Premodern Japanese Books LM Corpus 日本語 概要 日本古典籍統一データセットの言語モデル学習用本文ビュー v0.2.0 です。lm-curated v0.2.0から、本文採用対象とした1,598文書を収録しています。文書の本文はcontent列に入り、文書単位の論理分割はsplit列に記録しています。 収録範囲 kouigenji ndl-minhon-ocrdataset yatanavi 利用上の注意 Hugging Face上の物理splitはtrain一つです。split列にtrain、validation、testの論理分割を保持しています。 NDL Minhonでは角括弧の記号だけを削除し、角括弧内部の文字は保持しています。 やたナビでは読み仮名、異読、校訂注、その他の補助表記を除去しています。 利用条件と帰属表示はNOTICE.mdを確認してください。… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/premodern-japanese-books-lm-corpus.tabular1K<n<10K0 likes98 downloads19d agoHugging Face20llm-book /jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3 Dataset Card for "jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3" More Information needed tabular1M<n<10M1 likes83 downloads3y agoHugging Face21lynx1231 /historical-japanese-yen-crosses-sample Historical Japanese Yen Crosses Sample A free evaluation sample of historical Forex data for selected Japanese yen cross-currency pairs. Full historical Forex catalog, broader pair coverage, downloadable datasets, and pricing:https://futuresforexandsomeindexes.com/ This repository is a free evaluation sample intended for schema inspection, data-quality evaluation, integration testing, and quantitative research prototyping. It is not the complete commercial dataset.… See the full description on the dataset page: https://huggingface.co/datasets/lynx1231/historical-japanese-yen-crosses-sample.tabular1M<n<10M0 likes62 downloads8d agoHugging Face22Kotomiya07 /premodern-japanese-books-source-crosswalk Premodern Japanese Books Source Crosswalk 日本語 概要 日本古典籍統一データセット v0.1.0 の公開用Viewです。固定済み内部Releaseから、再配布と機械学習利用が許可された行だけを収録しています。収録行数は 17,570,138 行、収録SourceDataset数は 7 件です。 用途 日本語歴史資料の研究、検索、OCRまたは言語モデル用データ処理に利用できます。個々の行には採用したCanonical ID、権利判定、必要な帰属を保持しています。 権利 単一のライセンス値は全行の条件を表しません。必ず NOTICE.md と行単位の権利列を確認してください。 限界 v0.1.0 は監査対象52候補のうち、固定入力が成立した21 SourceDatasetを対象とする段階公開です。内容の正確性、外部参照の永続性、特定用途への適合性を保証しません。… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/premodern-japanese-books-source-crosswalk.tabular10M<n<100M0 likes60 downloads20d agoHugging Face23naive-puzzle /japanese-mt-benchtabularn<1K1 likes55 downloads2y agoHugging Face24p1atdev /japanese-stackexchange japanese-stackexchange 英語による日本語に関する質問ができる Japanese Stack Exchange のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。 日本語翻訳された StackExchange ではないです。 データ構造 投稿本文は html2text を使ってマークダウン化されています。その際、 コードブロックは ``` で囲まれるように変更されています。 画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。 default サブセット id: 質問投稿の ID question: 質問投稿 answers: 質問に対する回答投稿のリスト accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある simple サブセット… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/japanese-stackexchange.tabulartext-generation10K<n<100K3 likes54 downloads3y agoHugging Face25hotchpotch /japanese-qa-reasoning-100k 思考過程を含む、日本語質問・キーワード・回答・文章の合成データセット fineweb2-edu-japanese の文章データを元に、DeepSeek-R1 で文章(text)から質問文と回答部分の該当箇所を生成した日本語の質問と対応する文章・回答部のデータセットです。deekseek-r1 が出力した reasoning 部分も含まれます。testセットは、fineweb2-edu-japaneseのtestのみからサンプリングしています。 質問と文章ペアのデータセットやキーワードと文章ペアのデータセットとしてお使いいただけます。 ライセンス fineweb2 と同等の ODC-By とします。 tabular100K<n<1M3 likes47 downloads2y agoHugging Face26hotchpotch /japanese-query-crafter-reasoning-80k 思考過程を含む、クエリ作成のための日本語質問文テキストの合成データセット fineweb2-edu-japanese の small_tokens_cleaned の文章データを元に、DeepSeek-R1 で文章(text)から質問文を作成したデータセットです。deekseek-r1 が出力した reasoning 部分も含まれます。testセットは、fineweb2-edu-japaneseのtestのみからサンプリングしています。 ライセンス fineweb2 と同等の ODC-By とします。 tabular10K<n<100K3 likes45 downloads1y agoHugging Face27ronantakizawa /japanese-text-difficulty Aozora Text Difficulty Dataset This dataset contains Japanese literary texts from the Aozora Bunko digital library, enhanced with jReadability-based difficulty analysis for Japanese language learning and curriculum development. Dataset Overview Source: Aozora Bunko (青空文庫) - Japan's premier digital library of public domain literature Enhancement: jReadability-based difficulty scoring using research-backed Japanese readability models Primary Methodology: jReadability - A… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-text-difficulty.tabulartext-classification1K<n<10K5 likes42 downloads10mo agoHugging Face28argo11 /japanese-math-empirical-difficulty-pilot-50k Japanese Math Empirical Difficulty Pilot 50k This dataset is a 50,000-problem empirical difficulty pilot, not a full empirical labeling of the original 5.66M-row source dataset. It was created for LLM-jp experiment 0399, Team Victory SFT, to validate empirical difficulty label distribution, downstream split behavior, and the rollout/scoring pipeline before attempting labeling at the full 5.6M scale. Current Status This upload uses the v3 scorer with assistant-only… See the full description on the dataset page: https://huggingface.co/datasets/argo11/japanese-math-empirical-difficulty-pilot-50k.tabulartext-generation100K<n<1M0 likes42 downloads3mo agoHugging Face29Verah /JParaCrawl-Filtered-English-Japanese-Parallel-Corpus Introduction This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus. The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet. Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.tabulartranslation1M<n<10M3 likes37 downloads3y agoHugging Face30yzhuang /metatree_JapaneseVowels Dataset Card for "metatree_JapaneseVowels" More Information needed tabular1K<n<10K0 likes35 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.