datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imabari_wiki_qa_v4_validated
Imabari QA v4 — Validated
Dataset Summary
Imabari QA v4 — Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT).
This dataset combines two independently curated variants of the Imabari QA v4 synthetic reasoning dataset:
ikedachin/imabari_qa_v4_program_validated
ikedachin/imabari_qa_v4_human_validated
Both datasets are derived from:
Source corpus: ikedachin/imabari_wiki_cpt_v3
QA generation model: Qwen3.8-27B-NVFP4
The combined… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_validated.imabari_wiki_qa_v4_validated_w_reasoning_effort_qwen38
Imabari Wiki QA v4 Validated with Reasoning Effort — Qwen3.8
概要 / Overview
日本語・今治弁のQAを用いて、reasoning effort に応じた思考文の生成を学習するための教師ありファインチューニング(SFT)用データセットです。Imabari Wiki QA v4 Validated の質問と回答を保持し、元記事の文脈を参照して思考文を再生成しています。
This dataset supports supervised fine-tuning (SFT) of reasoning-effort-conditioned explanations using Japanese QA with Imabari dialect expressions. Questions and answers from Imabari Wiki QA v4 Validated are preserved, while reasoning text is… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_validated_w_reasoning_effort_qwen38.imabari_wiki_qa_v4_program_validated
Imabari QA v4 — Program Validated
Dataset Summary
Imabari QA v4 — Program Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT).
It is derived from:
Source corpus: ikedachin/imabari_wiki_cpt_v3
QA generation model: Qwen3.8-27B-NVFP4
The dataset contains synthetic question-answer pairs together with generated reasoning traces in the thinking field.
The primary difference from the original Imabari Wiki QA v4 dataset is the… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_program_validated.imabari_wiki_qa_v4_human_validated
Imabari QA v4 — Human Validated
Dataset Summary
Imabari QA v4 — Human Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT).
It is derived from:
Source corpus: ikedachin/imabari_wiki_cpt_v3
QA generation model: Qwen3.8-27B-NVFP4
The dataset contains synthetic question-answer pairs together with generated reasoning traces stored in the thinking field.
The primary characteristic of this dataset is that the generated samples have been… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_human_validated.std_lang_wiki_qa_v4_validated_w_reasoning_effort_llmjp4
Standard Japanese Wiki QA v4 with Reasoning Effort — LLM-jp 4
概要
Imabari Wiki QA v4 with Reasoning Effort — LLM-jp 4 の thinking(思考文)と answer(最終回答)の言葉遣いを標準語(です・ます調)へ変換した派生データセットです。日本語QAの教師ありファインチューニング(SFT)や、reasoning effort に応じた説明文・回答の生成に利用できます。
質問に新たに回答したり、推論内容を追加したりすることは目的としていません。元の意味・結論・根拠を保持し、今治弁などの方言表現を標準語へ置き換えるよう指示しています。質問は変換対象に含めていません。
LLM-jp 4 のチャットテンプレート用に、assistant の content に最終回答、thinking に思考文を格納しています。「LLM-jp 4」は対象のデータ形式を示します。元の思考文生成モデルと標準語化の設定モデルは… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/std_lang_wiki_qa_v4_validated_w_reasoning_effort_llmjp4.technical-docs-qa-validated
Technical Documentation Q&A - Validated
This is a validated version of nirav60614/technical-docs-qa with quality scores and filtering.
Validation Summary
Total Pairs: 261,077 (100%)
Valid Pairs: 248,096 (95.0%)
Average Quality Score: 0.867/1.0
Validation Method: LLM-based (llama3.2:latest via Ollama)
GPU: NVIDIA RTX 5090
Processing Time: ~28 hours
Validated: 2025-11-05
Quality Distribution
Quality Level
Score Range
Count
Percentage
Excellent
≥ 0.9… See the full description on the dataset page: https://huggingface.co/datasets/nirav60614/technical-docs-qa-validated.hausa-pq-speech-validated
Hausa WAEC PQ Speech Dataset (Validated)
Dataset Description
This dataset contains 1,361 multiple-choice WAEC past questions translated from English into Hausa, with human-validated Hausa translations. The data is designed to support speech synthesis, machine translation evaluation, and low-resource NLP research for Hausa — one of the most widely spoken languages in West Africa.
Languages
Source: English (en)
Target: Hausa (ha)… See the full description on the dataset page: https://huggingface.co/datasets/honourjesus/hausa-pq-speech-validated.
