CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llm-jp /leaderboard-requeststextn<1K2 likes18k downloads11mo agoHugging Face02Podtech /llm-jp-corpus-v4-ja_wiki llm-jp-corpus-v4 — ja_wiki Mirror of the ja/ja_wiki sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_wiki Files: 6 × jsonl.gz (1.9 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY-SA 3.0 — inherited from the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_wiki.texttext-generation1M<n<10M0 likes1k downloads2mo agoHugging Face03llm-jp /scaling-data-constrained-llms Scaling Data-Constrained Language Models with Synthetic Data This repository provides the pre-training corpora used in Scaling Data-Constrained Language Models with Synthetic Data (Findings of EACL 2026). Overview This repository contains multiple corpora designed to study data augmentation strategies for pre-training Japanese LLMs under a data-constrained data setting. Starting from a limited Japanese Web corpus and a larger English Web corpus, we construct three… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/scaling-data-constrained-llms.texttext-generation100M<n<1B5 likes822 downloads6mo agoHugging Face04Podtech /llm-jp-corpus-v4-ja_warp_pdf llm-jp-corpus-v4 — ja_warp_pdf Mirror of the ja/ja_warp_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_warp_pdf Files: 513 × jsonl.gz (73.8 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_pdf.texttext-generation10M<n<100M0 likes794 downloads2mo agoHugging Face05Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_html llm-jp-corpus-v4 — ja_sip_comprehensive_html Mirror of the ja/ja_sip_comprehensive_html sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_html Files: 181 × jsonl.gz (23.4 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_html.texttext-generation1M<n<10M0 likes646 downloads2mo agoHugging Face06Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_pdf llm-jp-corpus-v4 — ja_sip_comprehensive_pdf Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_pdf Files: 156 × jsonl.gz (39.1 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.texttext-generation1M<n<10M0 likes496 downloads2mo agoHugging Face07EeeLM /llm-jp-4-thinking-sft-data-chatmlllm-jpのデータセットllm-jp-4-thinking-sft-dataを、 ChatML形式に変換したものです。 ライセンス 各サンプルのライセンスは、元データセットカードに記載された各データソースのライセンスに従います。 本リポジトリは、元となったデータ全体に対して新たなライセンスを付与するものではありません。 利用する場合は、対応する元データソースのライセンス条件を確認してください。 text1M<n<10M0 likes479 downloads5mo agoHugging Face08Podtech /llm-jp-corpus-v4-ja_patent llm-jp-corpus-v4 — ja_patent Mirror of the ja/ja_patent sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_patent Files: 621 × jsonl.gz (58.2 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY 4.0 — inherited from… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_patent.texttext-generation1M<n<10M0 likes470 downloads2mo agoHugging Face09llm-jp /magpie-sft-v1.0 magpie-sft-v1.0 This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This is a dataset of instruction and response pairs created using the Magpie method. cyberagent/calm3-22b-chat was used for generating the instructions, and Qwen/Qwen2.5-32B-Instruct was used for generating the responses. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/magpie-sft-v1.0.texttext-generation100K<n<1M19 likes347 downloads2y agoHugging Face10Podtech /llm-jp-corpus-v4-ja_kokkai_giji llm-jp-corpus-v4 — ja_kokkai_giji Mirror of the ja/ja_kokkai_giji sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_kokkai_giji Files: 12 × jsonl.gz (1.3 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY 4.0 —… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_kokkai_giji.texttext-generation10K<n<100K0 likes270 downloads2mo agoHugging Face11llm-jp /oasst2-33k-ja oasst2-33k-ja This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan. The dataset comprises a Japanese translation of an English subset from oasst2, translated using DeepL. The English subset can be found here. For the creation of this dataset, we processed data from kunishou/oasst2-135k-ja. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/oasst2-33k-ja.text10K<n<100K13 likes253 downloads2y agoHugging Face12Podtech /llm-jp-corpus-v4-ja_warp_html llm-jp-corpus-v4 — ja_warp_html Mirror of the ja/ja_warp_html sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_warp_html Files: 45 × jsonl.gz (1.6 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_html.texttext-generation1M<n<10M0 likes212 downloads2mo agoHugging Face13llm-jp /mbpp-ja mbpp-ja This repository provides a mbpp dataset translated from English into Japanese by LLM-jp, a collaborative project launched in Japan. For English to Japanese translation, DeepL was used. The links of the original mbpp dataset are here(HuggingFace) or here(GitHub). Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order. Namgi Han, Masatoshi Otake, Shintaro Ozaki, Yusuke Miyao. textn<1K3 likes167 downloads2y agoHugging Face14llm-jp /oasst1-21k-ja oasst1-21k-ja This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This dataset is a Japanese translation of an English subset of oasst1 using DeepL. English subset is here. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order. Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/oasst1-21k-ja.text10K<n<100K17 likes138 downloads3y agoHugging Face15llm-jp /hh-rlhf-12k-ja hh-rlhf-12k-ja This repository provides a human preference dataset developed by LLM-jp, a collaborative project launched in Japan. This dataset is a Japanese translation of a subset of hh-rlhf using DeepL. This dataset consists of 12,000 entries randomly sampled from hh-rlhf. Specifically, it includes a random selection of 3,000 entries from the training splits of the four groups: harmless-base, helpful-base, helpful-online, and helpful-rejection-sampled. For more information on… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/hh-rlhf-12k-ja.text10K<n<100K15 likes136 downloads3y agoHugging Face16llm-jp /databricks-dolly-15k-ja databricks-dolly-15k-ja This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This dataset is a Japanese translation of databricks-dolly-15k using DeepL. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order. Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi Nakamura, Takashi Kodama, Takumi… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/databricks-dolly-15k-ja.textquestion-answering10K<n<100K18 likes118 downloads3y agoHugging Face17Podtech /llm-jp-corpus-v4-ja_kaken llm-jp-corpus-v4 — ja_kaken Mirror of the ja/ja_kaken sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_kaken Files: 5 × jsonl.gz (1.1 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY 4.0 — inherited from the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_kaken.texttext-generation1M<n<10M0 likes84 downloads2mo agoHugging Face18Podtech /llm-jp-corpus-v4-ja_e-gov llm-jp-corpus-v4 — ja_e-gov Mirror of the ja/ja_e-gov sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_e-gov Files: 2 × jsonl.gz (0.1 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY 4.0 — inherited from the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_e-gov.texttext-generation1K<n<10K0 likes68 downloads2mo agoHugging Face19p1atdev /LLM-jp-Toxicity-Dataset LLM-jp Toxicity Dataset 日本語有害文書データセット「LLM-jp Toxicity Dataset」 See https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-toxicity-dataset texttext-classification1K<n<10K5 likes64 downloads2y agoHugging Face20AcroYAMALEX /acro-yamalex-llmjp-4-math-tir acro-yamalex-llmjp-4-math-tir(データセット) 日本語数学推論のためのTool-Integrated Reasoning (TIR) データセットです。 自然言語による推論とPythonコード実行を組み合わせたマルチターン形式のデータセットで、OpenWebMathから抽出・生成した問題に対してDeepSeek V3を用いてTIR形式の解法を生成しました。 本データセットはFT-LLM2026コンペティションにおける我々のアプローチの一部として構築されました。 データセット概要 項目 値 データ件数 134,834件 言語 日本語 ソース OpenWebMath 生成モデル DeepSeek V3 (deepseek-chat) フォーマット JSONL(マルチターン対話形式) データ作成手法 OpenMathReasoning(NVIDIAのAIMO-2優勝手法)の枠組みに基づき、以下の手順で作成しました。… See the full description on the dataset page: https://huggingface.co/datasets/AcroYAMALEX/acro-yamalex-llmjp-4-math-tir.texttext-generation100K<n<1M0 likes51 downloads6mo agoHugging Face21llm-jp /JSFactCheckBenchgated JSFactCheckBench 概要 JSFactCheckBenchは、LLM応答のファクトチェックを評価するための日本語ベンチマークです。 日本のソーシャルメディア上で流通した誤情報に由来する85問の質問(JSocialFactより)に対して、11 種類の LLM が生成した応答 170 件をベースに、1,933 件の言明を作成しました。 各言明には、(1) 事実性ラベル(Supported / Refuted / NotEnoughInfo、補助ラベル Unevaluable / Unclear)と、(2) 質問に対する役割を表す重要度ラベル(Main / Sub / Irrelevant)の2軸を人手で付与しています。 事実性は、根拠となるWebページを収集した時点を基準時刻として、その根拠に基づいて判定しています。 詳細は、後日公開予定の論文をご覧ください。 Overview JSFactCheckBench is a Japanese benchmark for evaluating the fact-checking of LLM… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/JSFactCheckBench.texttext-classification1K<n<10K0 likes51 downloads5d agoHugging Face22AcroYAMALEX /acro-yamalex-llmjp-4-math-cot acro-yamalex-llmjp-4-math-cot(データセット) 日本語数学推論のためのChain-of-Thought (CoT) データセットです。 StackMathQAの問題に対して、DeepSeek V3を用いて日本語CoT形式の解法を生成しました。 本データセットはFT-LLM2026コンペティションにおける我々のアプローチの一部として構築されました。 データセット概要 項目 値 データ件数 306,366件 言語 日本語 ソース StackMathQA 生成モデル DeepSeek V3 (deepseek-chat) フォーマット JSONL(マルチターン対話形式) データ作成手法 OpenMathReasoning(NVIDIAのAIMO-2優勝手法)の枠組みに基づき、以下の手順で作成しました。 Step 1: 日本語CoT解法の生成 StackMathQAの問題に対してDeepSeek… See the full description on the dataset page: https://huggingface.co/datasets/AcroYAMALEX/acro-yamalex-llmjp-4-math-cot.texttext-generation100K<n<1M1 likes50 downloads6mo agoHugging Face23llm-jp /oasst1-21k-en oasst1-21k-en This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This dataset is an English subset of oasst1. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order. Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi Nakamura, Takashi Kodama, Takumi Okamoto. text10K<n<100K2 likes40 downloads3y agoHugging Face24llm-jp /llava-instruct-ja Dataset Card for llava_instruct_ja Dataset details This is the Japanese version of LLaVA-Instruct, which contains 156K samples. We used gpt-4o-mini-2024-07-18 to generate data through via Azure OpenAI API. License Creative Commons Attribution 4.0 License; and it should abide by the OpenAI terms of use textvisual-question-answering100K<n<1M5 likes40 downloads2y agoHugging Face25Podtech /llm-jp-corpus-v4-ja_wiki-manufacturing ja_wiki 製造業コーパス llm-jp-corpus-v4 の ja/ja_wiki(120 万記事)から、製造業に関連する文書を抽出した継続事前学習(CPT)用データ。 文書数 47,914 件(元コーパスの 3.99%) 文字数 105M 字(推定 74M トークン) 形式 JSONL・1 行 1 文書 抽出方法 キーワード一致では役に立たない。「製造」「工場」で本文先頭を引くと 5.7% が命中するが、 中身は工場や製造に一言触れただけの記事が大半だった。そこで 精度を Wikipedia の カテゴリ木、網羅性を分類器に担わせ、その和集合を採っている。 カテゴリ木 — 製造業・業種別・機械工学・材料工学・産業遺産など 42 のシードから 日本語 Wikipedia のカテゴリグラフを深さ 3 まで辿る(SQL ダンプからオフライン構築)。 突合 — 得られたタイトルを ja_wiki の実在記事に絞る。 分類器 — 上記を正例、無作為抽出を負例として、文字… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_wiki-manufacturing.texttext-generation10K<n<100K0 likes37 downloads2mo agoHugging Face26llm-jp /aya-ja-evol-inst aya-ja-eval-inst This repository provides a preference dataset developed by LLM-jp, a collaborative project launched in Japan. This dataset was created by generating the chosen response using Qwen/Qwen2.5-32B-Instruct and the rejected response using llm-jp/llm-jp-3-1.8b-instruct for the prompts in weblab-GENIAC/aya-ja-evol-instruct-calm3-dpo-masked. This repository does not contain prompts but only the corresponding indices. Please obtain the original prompts from the original data… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/aya-ja-evol-inst.texttext-generation10K<n<100K1 likes36 downloads2y agoHugging Face27llm-jp /ja-vg-vqa-conversation Dataset Card for ja-vg-vqa-conversation Dataset details This dataset consists of multiple QA pairs from Japanese Visual Genome VQA dataset that have been converted into multi-turn conversations. The prompt 語句または短い文で答えてください。 has been appended to the questions in the first turn. We excluded the samples from the JA-VG-VQA-500 dataset, resulting in a total of 98,708 samples. License Creative Commons Attribution 4.0 License textvisual-question-answering10K<n<100K3 likes34 downloads2y agoHugging Face28llm-jp /japanese-photos-conversation Dataset Card for japanese photos conversation Dataset details This dataset contains multi-turn conversational instructions about images taken in Japan. The images were sourced from https://huggingface.co/datasets/ThePioneer/japanese-photos. We input each image into GPT-4o (gpt-4o-2024-05-13) via the Azure OpenAI API to generate the instruction data. Some of the images in the original dataset were filtered by the Azure OpenAI API when they were input, resulting in a total… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/japanese-photos-conversation.textvisual-question-answering10K<n<100K7 likes24 downloads2y agoHugging Face29open-llm-leaderboard /jpacifico__Chocolatine-14B-Instruct-4k-DPO-detailsgated Dataset Card for Evaluation run of jpacifico/Chocolatine-14B-Instruct-4k-DPO Dataset automatically created during the evaluation run of model jpacifico/Chocolatine-14B-Instruct-4k-DPO The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jpacifico__Chocolatine-14B-Instruct-4k-DPO-details.tabular10K<n<100K0 likes20 downloads2y agoHugging Face30DeL-TaiseiOzaki /magpie-llm-jp-3-13b-20k 合成日本語指示データセット 概要 このデータセットは、大規模言語モデル(LLM)を用いて自動生成された日本語の指示とそれに対する応答のコレクションです。データセットは指示付与型のタスクのための学習や評価に使用することを目的としています。 データセット仕様 サンプル数: 20,000 言語: 日本語 フォーマット: JSON 生成方法 データセットは以下のプロセスを通じて生成されました: LLM-JP 3.13B Instructモデルを使用 各サンプルは3段階のプロセスで生成: a) 指示文の生成 b) Chain-of-Thought (CoT) 応答の生成 (一部のデータには含まれない) c) 最終的な応答のself-refine https://github.com/DeL-TaiseiOzaki/magpie-llm-jp-3 データ構造 各サンプルは以下の構造を持つJSONオブジェクトです: { "instruction": "指示文"… See the full description on the dataset page: https://huggingface.co/datasets/DeL-TaiseiOzaki/magpie-llm-jp-3-13b-20k.text10K<n<100K0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.