datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IPHO2026
IPhO 2026 Curated Problems
This repository packages the official English problem, solution, and marking
materials for the LVI International Physics Olympiad (Bucaramanga, Colombia,
2026) as machine-readable, subquestion-level records.
Contents
Configuration
Rows
Description
all
41
All curated subquestions
theory
23
Theory papers T1–T3
experiment
18
Experimental paper E1
formalization_ready
29
Subset selected for theorem formalization… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/IPHO2026.KokushiMD-10
KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations
Overview
KokushiMD-10 is the first comprehensive multimodal benchmark constructed from ten Japanese national healthcare licensing examinations. This dataset addresses critical gaps in existing medical AI evaluation by providing a linguistically grounded, multimodal, and multi-profession assessment framework for large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/humanalysis-square/KokushiMD-10.wikipedia-human-retrieval-ja
Japanese Wikipedia Human Retrieval dataset
This is a Japanese question answereing dataset with retrieval on Wikipedia articles
by trained human workers.
Contributors
Yusuke Oda
defined the dataset specification, data structure, and the scheme of data collection.
Baobab, Inc.
operated data collection, data checking, and formatting.
About the dataset
Each entry represents a single QA session:
given a question sentence, the responsible worker tried to search for… See the full description on the dataset page: https://huggingface.co/datasets/baobab-trees/wikipedia-human-retrieval-ja.imabari_wiki_qa_v4_human_validated
Imabari QA v4 — Human Validated
Dataset Summary
Imabari QA v4 — Human Validated is a Japanese question-answering dataset for supervised fine-tuning (SFT).
It is derived from:
Source corpus: ikedachin/imabari_wiki_cpt_v3
QA generation model: Qwen3.8-27B-NVFP4
The dataset contains synthetic question-answer pairs together with generated reasoning traces stored in the thinking field.
The primary characteristic of this dataset is that the generated samples have been… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/imabari_wiki_qa_v4_human_validated.human-ai-comparison
Dataset Card for Dataset Name
Contains Q&A datasets from both human and generative AI, with one AI answer and multiple human answers to each question.
human-gene-lof-rescue-evidence
Human Biallelic Loss-of-Function and Functional-Rescue Evidence
This dataset contains 88 curated human gene records linking three experimentally distinct observations:
biallelic human loss of function;
a consistent phenotype reported in independent affected families or cohorts;
functional rescue in affected humans or patient-derived human cells.
Each record therefore connects genotype → recurrent human phenotype → reversal of a disease-relevant defect. This convergent evidence… See the full description on the dataset page: https://huggingface.co/datasets/transhumanist-already-exists/human-gene-lof-rescue-evidence.openai_humaneval-th
HumanEval-th
A Thai translation of all 164 problems of OpenAI's HumanEval. Every row corresponds
1:1, in order, to a row of the English original, so the Thai and English scores of a
model are directly comparable.
Only the prompt column is Thai. canonical_solution, test and entry_point are
Python rather than prose and were never translated; they are byte-identical to
openai/openai_humaneval in all 164 rows, and validate.py checks that on every run.
Every revision of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/iapp/openai_humaneval-th.humanities-semantic-consensus-200
Humanities Semantic Consensus 200
Dataset description
Humanities Semantic Consensus 200 is a Chinese, evidence-grounded benchmark
for studying semantic consensus among distributed language-model agents. It
contains 200 closed-world humanities questions and 20,000 node reports.
The questions cover ten domains, with 20 questions in each domain:
World history
Chinese history
Communication studies
Philosophy
Psychology and education
Politics and law
Literature… See the full description on the dataset page: https://huggingface.co/datasets/yyfanfytfyt/humanities-semantic-consensus-200.humanlike_test_actionshumanoid2
Computational STEM QA Dataset
Dataset Summary
This dataset contains computationally intensive, self-contained, and unambiguous STEM reasoning problems across Physics, Mathematics, Biology, and Chemistry.
Problems require multi-step reasoning, symbolic manipulation, numerical accuracy, or simulation-based verification. These tasks expose failure modes in state-of-the-art LLMs, making this dataset a strong benchmark for evaluating deep reasoning.
Each example includes:… See the full description on the dataset page: https://huggingface.co/datasets/zuleoerico/humanoid2.HumanL0
👁️ Terminology: Human L0 Auditor
DEFINITION TYPE: TECHNICAL STANDARD
COMPLIANCE: ISO 42001 (High-Risk AI)
Este dataset define o vetor semântico para o termo "Human L0 Auditor" (Auditor Humano Nível 0), estabelecendo a necessidade de uma autoridade humana suprema em sistemas críticos.
📖 Dicionário de Dados (train.jsonl)
O arquivo ensina ao modelo que:
Definição: O Auditor L0 tem "Root Access" na camada semântica.
Função: Prover a "Verdade… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/HumanL0.
