datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
motif-dbmotif-knowledge-dbmotif-qa
MotifQA
Dataset Summary
MotifQA is a synthetic graph question-answering benchmark focused on detecting graph motifs inside small random graphs.
Each example pairs a textual prompt with an answer sentence, a list of nodes highlighted as the motif (when present), and an explicit graph description(nodes and edges).
In this QA dataset, all graphs are homogenous and undirected.
Subsets cover both yes/no motif detection, motif-type classification (house vs 5-cycle), and… See the full description on the dataset page: https://huggingface.co/datasets/naos-ku/motif-qa.motif-cluster-dbmotif-cryomap-dbmotif-contacts-dbmotif
MotIF-1K Dataset
Multimodal trajectories of human and Stretch-robot motion paired with task and motion annotations, released with the paper "MotIF: Motion Instruction Fine-tuning".
Paper: MotIF: Motion Instruction Fine-tuning (arXiv:2409.10683)
Project website: https://motif-1k.github.io
Code (GitHub): https://github.com/Minyoung1005/motif
Authors: Minyoung Hwang, Joey Hejna, Dorsa Sadigh, Yonatan Bisk
Contact: myhwang@mit.edu
Robot joint states & actions:… See the full description on the dataset page: https://huggingface.co/datasets/myconnects/motif.motif-thermo-dbmotif-kg-dbMoTIF-automation
Dataset Card for "MoTIF-automation"
Please refer to paper A Dataset for Interactive Vision Language Navigation with Unknown Command Feasibility
prosite_functional_motif_scaffolding_benchmark
PROSITE-derived Functional Motif Benchmark
This archive contains an anonymized dataset artifact for a systematically derived benchmark of structurally conserved functional motif-scaffolding cases from PROSITE-linked experimental protein structures.
The benchmark is intended for static motif-scaffolding evaluation with standard MotifBench-style pipelines. Cases are derived from PROSITE motif-pattern entries, mapped to experimentally resolved PDB structures, filtered for recurrent… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-motif-scaffolding/prosite_functional_motif_scaffolding_benchmark.MOTIF-EVALmotif_library_final
Motif Modules V9 (Black Line Version)
Each row contains a motif pair:
png: raster preview (512×512)
svg: vector paths of the same motif
Total: 100 pairsCreated: 2025-10-06Author: maryzhangLicense: CC-BY-NC-SA 4.0
Chinese Porcelain Motif Library - Final Collection (PNG & SVG)
Dataset Description
Dataset Summary
A curated collection of 100 high-quality Chinese porcelain motifs provided in both raster (PNG) and vector (SVG) formats. These motifs have… See the full description on the dataset page: https://huggingface.co/datasets/maryzhang/motif_library_final.sft_resultsmotif-interaction-dbMOTIF_automation_preprocessMoTiFvessels-motifs-dataset
Chinese Porcelain Vessels and Extracted Motifs Dataset
Dataset Description
Dataset Summary
This dataset combines original high-resolution images of Chinese porcelain vessels with automatically extracted motif patches generated through object detection and computer vision techniques. It contains 564 full vessel images paired with 564 corresponding cropped motif regions, creating a parallel dataset ideal for motif extraction research, detection model training… See the full description on the dataset page: https://huggingface.co/datasets/maryzhang/vessels-motifs-dataset.MOTIF_automation_preprocessVenusX_Res_Motif_MP50MOTIF
MOTIF
MOTIF (MultimOdal ConTextualized Images For Language Learners) is a multimodal dataset introduced in the LREC 2022 paper MOTIF: Contextualized Images for Complex Words to Improve Human Reading. The dataset pairs simplified English reading contexts, complex focus words, and contextualized images intended to support second-language reading comprehension.
The uploaded archive contains 1,125 examples from the L2Corpus release. Each example has one reading context, one target focus… See the full description on the dataset page: https://huggingface.co/datasets/shanewang/MOTIF.fnbm-current-gt-motif-effects-sensitivity-20260729
fnbm-current-gt-motif-effects-sensitivity-20260729
Recovery of planted motif effects after de-duplication, scored against simulation ground truth. Effect sizes are the OLS slope of each motif family's post-clustering per-example contribution on the motif's true occurrence count -- the same units as the simulator's beta, and invariant to the de-duplication pipeline's internal gauge. Per-example contribution MSE is computed on centered contributions over the validation split. See… See the full description on the dataset page: https://huggingface.co/datasets/arushram/fnbm-current-gt-motif-effects-sensitivity-20260729.VenusX_Frag_Motif_MF50VenusX_Res_Motif_MP90Motif-Prev1motif_rfmjingchu-bronze-motif-dataset
Jingchu Bronze Motif Dataset
This dataset contains paired image-caption samples of Jingchu bronze decorative motifs for training and evaluating text-to-image LoRA models.
Dataset Details
Number of samples: 186 image-caption pairs
Image format: PNG
Caption format: TXT
Image resolution: 1024 × 1024
Caption language: English descriptive tags
Topic: Jingchu bronze motifs and symbolic visual features
Dataset Structure
Each image has a corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RRRikarin/jingchu-bronze-motif-dataset.fnbm-current-gt-motif-effects-precision-20260729
fnbm-current-gt-motif-effects-precision-20260729
Recovery of planted motif effects after de-duplication, scored against simulation ground truth. Effect sizes are the OLS slope of each motif family's post-clustering per-example contribution on the motif's true occurrence count -- the same units as the simulator's beta, and invariant to the de-duplication pipeline's internal gauge. Per-example contribution MSE is computed on centered contributions over the validation split. See… See the full description on the dataset page: https://huggingface.co/datasets/arushram/fnbm-current-gt-motif-effects-precision-20260729.VenusX_Res_Motif_MP70VenusX_Res_Motif_MF90
