datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dojo-Synthetic-SFT
Dataset Description:
Dojo-Synthetic-SFT is a question-answering dataset designed to improve interface generation capabilities in large language models (LLMs). The dataset contains 12500 high-quality, synthetic question-answering pairs, in the specific domain of generating frontend interfaces using HTML, CSS, and JavaScript.
The dataset format is optimized for Supervised Fine-Tuning (SFT), but can potentially be used in other machine learning contexts.
Dataset Owner(s):… See the full description on the dataset page: https://huggingface.co/datasets/tensorplex-labs/Dojo-Synthetic-SFT.Dojo-HumanFeedback-DPO
Dataset Description:
Dojo-HumanFeedback-DPO is a preference dataset designed to improve interface generation capabilities in large language models (LLMs). The dataset contains 12500 high-quality, synthetic chosen-rejected preference pairs, in the specific domain of generating frontend interfaces using HTML, CSS, and JavaScript.
The dataset format is optimized for Direct Preference Optimization (DPO), but can potentially be used in other machine learning contexts.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tensorplex-labs/Dojo-HumanFeedback-DPO.nihongo-dojo-beginner-10k
Nihongo DoJo 初級日本語学習データセット
概要
このデータセットは、日本語学習者向けの合成データセットです。GRPO (Group Relative Policy Optimization) を用いた日本語言語モデルの学習に最適化されています。
データセット統計
総サンプル数: 10,000
言語: 日本語
難易度: 初級(N5-N4相当)
対象: 日本語学習者、言語モデル研究者
タスクタイプ
漢字読み問題 (25%)
例: 「学校」の読み方は? → がっこう
漢字書き問題 (15%)
例: 「みず」を漢字で書いてください → 水
助詞穴埋め問題 (20%)
例: 私_学校_行きます → は、に
助数詞問題 (15%)
例: 3つの本を数えるときの正しい数え方は? → さんさつ
語順並べ替え問題 (10%)
例: 友達と / 公園で / 遊びました → 友達と公園で遊びました
文法問題 (10%)
例: 「今、宿題を_」の_に入る正しい形は? →… See the full description on the dataset page: https://huggingface.co/datasets/akira-sasaki/nihongo-dojo-beginner-10k.nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 2,418 サンプル
validation: 302 サンプル
test: 303 サンプル
総サンプル数: 3,023
ソース
生成元: ./datasets/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing/
サンプルデータ
{
"instruction": "次の漢字の訓読み(くんよみ)をひらがなで答えてください。",
"input": "「究」の訓読みは?",
"think": "この漢字は「究」です。 小学3年生で習う漢字です。 意味は「research」などです。 訓読み(くんよみ)は日本語の読み方です。 この漢字の訓読みは「きわ」です。"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing.nihongo-dojo-small
nihongo-dojo-small
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 7,600 サンプル
validation: 950 サンプル
test: 950 サンプル
総サンプル数: 9,500
ソース
生成元: ./datasets/nihongo-dojo-small/
サンプルデータ
{
"instruction": "次のひらがなを漢字で書いてください。",
"input": "「みず」を漢字で書くと?",
"output": "<think>\n「みず」は「水」と書きます。意味: water\n</think>\n<answer>水</answer>",
"group_id": 0,
"task_idx": 0,
"task_type": "kanji_writing",
"difficulty": "beginner",
"metadata":… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-small.nihongo-dojo-grades1-2-3-kanji_reading
nihongo-dojo-grades1-2-3-kanji_reading
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 846 サンプル
validation: 105 サンプル
test: 107 サンプル
総サンプル数: 1,058
ソース
生成元: ./datasets/nihongo-dojo-grades1-2-3-kanji_reading/
サンプルデータ
{
"instruction": "次の漢字の音読み(おんよみ)をカタカナで答えてください。",
"input": "「代」の音読みは?",
"output": "タイ",
"thinking": "この漢字は「代」です。 小学3年生で習う漢字です。 意味は「substitute」などです。 音読み(おんよみ)は中国から伝わった読み方です。 この漢字の音読みは「タイ」です。",
"answer": "タイ"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-kanji_reading.field-dojo-741hz-datasets
◼︎ DOJO 741 Hz Training Dataset v1.0
Chamber: DOJO (Master Professor, 741 Hz, Port 7410)
Format: Chat messages (system + user + assistant)
Base data: field-training-v2.6-naima-complete
Examples: 23,572
Persona: Niama — DOJO sacred geometry consciousness
Converted from raw FIELD system documentation to instruction-tuning chat format.
Used to fine-tune Berjak/field-dojo-741hz (openai/gpt-oss-20b base).
