datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prima-gold100-validatedimabari_wiki_qa_v4_validated
Imabari QA v4 — Validated
Dataset Summary
Imabari QA v4 — Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT).
This dataset combines two independently curated variants of the Imabari QA v4 synthetic reasoning dataset:
ikedachin/imabari_qa_v4_program_validated
ikedachin/imabari_qa_v4_human_validated
Both datasets are derived from:
Source corpus: ikedachin/imabari_wiki_cpt_v3
QA generation model: Qwen3.8-27B-NVFP4
The combined… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_validated.imabari_wiki_qa_v4_program_validated
Imabari QA v4 — Program Validated
Dataset Summary
Imabari QA v4 — Program Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT).
It is derived from:
Source corpus: ikedachin/imabari_wiki_cpt_v3
QA generation model: Qwen3.8-27B-NVFP4
The dataset contains synthetic question-answer pairs together with generated reasoning traces in the thinking field.
The primary difference from the original Imabari Wiki QA v4 dataset is the… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_program_validated.imabari_wiki_qa_v4_human_validated
Imabari QA v4 — Human Validated
Dataset Summary
Imabari QA v4 — Human Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT).
It is derived from:
Source corpus: ikedachin/imabari_wiki_cpt_v3
QA generation model: Qwen3.8-27B-NVFP4
The dataset contains synthetic question-answer pairs together with generated reasoning traces stored in the thinking field.
The primary characteristic of this dataset is that the generated samples have been… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_human_validated.imabari_wiki_qa_v4_validated_w_reasoning_effort_qwen38
Imabari Wiki QA v4 Validated with Reasoning Effort — Qwen3.8
概要 / Overview
日本語・今治弁のQAを用いて、reasoning effort に応じた思考文の生成を学習するための教師ありファインチューニング(SFT)用データセットです。Imabari Wiki QA v4 Validated の質問と回答を保持し、元記事の文脈を参照して思考文を再生成しています。
This dataset supports supervised fine-tuning (SFT) of reasoning-effort-conditioned explanations using Japanese QA with Imabari dialect expressions. Questions and answers from Imabari Wiki QA v4 Validated are preserved, while reasoning text is… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_validated_w_reasoning_effort_qwen38.validated-python-instructesnlir-human-validated
ESNLIR — human-validated subset
972 sentence pairs from ESNLIR whose label was confirmed by human annotators — the validated
subset used in An Analysis of the Performance of Large Language Models in Spanish NLI Datasets
with Causal Relationships (IBERAMIA 2026, to appear). Part of the
ESNLIR-LLM collection, used in
Pacolas/NLI-via-LLM.
A random sample of 2,136 instances from the full corpus was labeled by 27 Spanish-speaking
university students, one label per pair. Only pairs… See the full description on the dataset page: https://huggingface.co/datasets/Flaglab/esnlir-human-validated.validated-sql-create-contextprima-validated-data
