datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
motif-qa
MotifQA
Dataset Summary
MotifQA is a synthetic graph question-answering benchmark focused on detecting graph motifs inside small random graphs.
Each example pairs a textual prompt with an answer sentence, a list of nodes highlighted as the motif (when present), and an explicit graph description(nodes and edges).
In this QA dataset, all graphs are homogenous and undirected.
Subsets cover both yes/no motif detection, motif-type classification (house vs 5-cycle), and… See the full description on the dataset page: https://huggingface.co/datasets/naos-ku/motif-qa.motif-cluster-dbmotif-contacts-dbmotif-cryomap-dbmotif-thermo-dbmotif-kg-dbprosite_functional_motif_scaffolding_benchmark
PROSITE-derived Functional Motif Benchmark
This archive contains an anonymized dataset artifact for a systematically derived benchmark of structurally conserved functional motif-scaffolding cases from PROSITE-linked experimental protein structures.
The benchmark is intended for static motif-scaffolding evaluation with standard MotifBench-style pipelines. Cases are derived from PROSITE motif-pattern entries, mapped to experimentally resolved PDB structures, filtered for recurrent… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-motif-scaffolding/prosite_functional_motif_scaffolding_benchmark.MOTIF_automation_preprocessmotif-interaction-dbMOTIF_automation_preprocessfnbm-current-gt-motif-effects-sensitivity-20260729
fnbm-current-gt-motif-effects-sensitivity-20260729
Recovery of planted motif effects after de-duplication, scored against simulation ground truth. Effect sizes are the OLS slope of each motif family's post-clustering per-example contribution on the motif's true occurrence count -- the same units as the simulator's beta, and invariant to the de-duplication pipeline's internal gauge. Per-example contribution MSE is computed on centered contributions over the validation split. See… See the full description on the dataset page: https://huggingface.co/datasets/arushram/fnbm-current-gt-motif-effects-sensitivity-20260729.fnbm-current-gt-motif-effects-precision-20260729
fnbm-current-gt-motif-effects-precision-20260729
Recovery of planted motif effects after de-duplication, scored against simulation ground truth. Effect sizes are the OLS slope of each motif family's post-clustering per-example contribution on the motif's true occurrence count -- the same units as the simulator's beta, and invariant to the de-duplication pipeline's internal gauge. Per-example contribution MSE is computed on centered contributions over the validation split. See… See the full description on the dataset page: https://huggingface.co/datasets/arushram/fnbm-current-gt-motif-effects-precision-20260729.detergent-motifbb_motifs_by_strandaxisba_motifsbb_motifs_localab_motifsmaqam-labels-rawspatial-mmcot-motif
Spatial MMCoT v1 · motif
MoTiF (OpenRaiser/MoTiF), the naive arm of its procedurally generated multi-step tasks: ball_tracking_naive, manipulation_naive, maze_naive (manipulation shows MoTiF's own CLEVR-style 3D renders of solids, not images from the CLEVR dataset). Each step has its own target image. The reflexion arm, which pairs each problem with a deliberately corrupted frame, is not included. sokoban_naive is excluded (7,997 upstream rows, S0.task_excluded): its eight… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-motif.
