datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
motif-1k
Dataset Card for MotIF-1K
MotIF-1K is a robotics motion dataset containing 1,022 demonstrations across 13 task categories, used to benchmark and fine-tune vision-language models (VLMs) for motion-based success detection. Each demonstration includes a video of the motion, multiple pre-rendered trajectory visualizations, task instructions, and motion descriptions.
The FiftyOne dataset is a grouped dataset where each group represents one trajectory and each group slice represents a… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/motif-1k.motif-dbmotif-knowledge-dbmotif-qa
MotifQA
Dataset Summary
MotifQA is a synthetic graph question-answering benchmark focused on detecting graph motifs inside small random graphs.
Each example pairs a textual prompt with an answer sentence, a list of nodes highlighted as the motif (when present), and an explicit graph description(nodes and edges).
In this QA dataset, all graphs are homogenous and undirected.
Subsets cover both yes/no motif detection, motif-type classification (house vs 5-cycle), and… See the full description on the dataset page: https://huggingface.co/datasets/naos-ku/motif-qa.motif-cluster-dbmotif-contacts-dbmotif-cryomap-dbmotif
MotIF-1K Dataset
Multimodal trajectories of human and Stretch-robot motion paired with task and motion annotations, released with the paper "MotIF: Motion Instruction Fine-tuning".
Paper: MotIF: Motion Instruction Fine-tuning (arXiv:2409.10683)
Project website: https://motif-1k.github.io
Code (GitHub): https://github.com/Minyoung1005/motif
Authors: Minyoung Hwang, Joey Hejna, Dorsa Sadigh, Yonatan Bisk
Contact: myhwang@mit.edu
Robot joint states & actions:… See the full description on the dataset page: https://huggingface.co/datasets/myconnects/motif.motif-thermo-dbMOSAIC_conditional_motif
MOSAIC: Conditional Motif Generation
Conditional generation samples + metrics for the three MOSAIC tokenizers
(SENT, HDT-MC, HDTC) prompted with each of 6 shared ring/aromatic motifs,
across MOSES, GuacaMol, and COCONUT.
The conditioning prompt for each model is the standalone-motif tokenization,
with closing tokens stripped so the model continues from "the molecule starts
with this motif." For HDT-MC, the outer ENTER block is left open so the model
can add additional communities;… See the full description on the dataset page: https://huggingface.co/datasets/MOSAIC-UCSD/MOSAIC_conditional_motif.MoTIF-automation
Dataset Card for "MoTIF-automation"
Please refer to paper A Dataset for Interactive Vision Language Navigation with Unknown Command Feasibility
motif-kg-dbMOTIF-EVALprosite_functional_motif_scaffolding_benchmark
PROSITE-derived Functional Motif Benchmark
This archive contains an anonymized dataset artifact for a systematically derived benchmark of structurally conserved functional motif-scaffolding cases from PROSITE-linked experimental protein structures.
The benchmark is intended for static motif-scaffolding evaluation with standard MotifBench-style pipelines. Cases are derived from PROSITE motif-pattern entries, mapped to experimentally resolved PDB structures, filtered for recurrent… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-motif-scaffolding/prosite_functional_motif_scaffolding_benchmark.MoTiFmotif_actions
MotIF-Actions
Robot action-space companion to the MotIF-1K
dataset. The public MotIF release
(myconnects/motif) exposes only the 2D end-effector pixel
trajectory; MotIF-Actions adds the on-board Stretch proprioception for the
subset of demonstrations that were logged with joint states.
Built directly from the raw teleoperation recordings, in
RLDS / TFDS format.
For the full MotIF-1K release — videos, trajectory overlays, storyboards, and
optical flow for both human and robot… See the full description on the dataset page: https://huggingface.co/datasets/myconnects/motif_actions.MOTIF_automation_preprocessBeautiful-Motifs
Beautiful Motifs
1.06M+ highest-rated beautiful music motifs extracted from 2.09M+ select HQ MIDIs from the Discover MIDI dataset
Dataset information
Reference MIDI dataset: Beautiful Music Seeds
Rating method: midisim
Beauty metric: top k embeddings cosine similarity against transposed reference MIDI dataset
Beauty scores range (min, mean, max): 0.7099544, 0.9020204, 0.9999562
Beauty rating information
Detailed beauty rating… See the full description on the dataset page: https://huggingface.co/datasets/asigalov61/Beautiful-Motifs.motif-interaction-dbmotif-image-galleryMOTIF_automation_preprocesssft_resultsVenusX_Res_Motif_MP50sft_results_2unbalanced-motifs-500K
Dataset Card for "unbalanced-motifs-500K"
More Information needed
jingchu-bronze-motif-dataset
Jingchu Bronze Motif Dataset
This dataset contains paired image-caption samples of Jingchu bronze decorative motifs for training and evaluating text-to-image LoRA models.
Dataset Details
Number of samples: 186 image-caption pairs
Image format: PNG
Caption format: TXT
Image resolution: 1024 × 1024
Caption language: English descriptive tags
Topic: Jingchu bronze motifs and symbolic visual features
Dataset Structure
Each image has a corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RRRikarin/jingchu-bronze-motif-dataset.MOTIF
MOTIF
MOTIF (MultimOdal ConTextualized Images For Language Learners) is a multimodal dataset introduced in the LREC 2022 paper MOTIF: Contextualized Images for Complex Words to Improve Human Reading. The dataset pairs simplified English reading contexts, complex focus words, and contextualized images intended to support second-language reading comprehension.
The uploaded archive contains 1,125 examples from the L2Corpus release. Each example has one reading context, one target focus… See the full description on the dataset page: https://huggingface.co/datasets/shanewang/MOTIF.VenusX_Frag_Motif_MF50fnbm-current-gt-motif-effects-sensitivity-20260729
fnbm-current-gt-motif-effects-sensitivity-20260729
Recovery of planted motif effects after de-duplication, scored against simulation ground truth. Effect sizes are the OLS slope of each motif family's post-clustering per-example contribution on the motif's true occurrence count -- the same units as the simulator's beta, and invariant to the de-duplication pipeline's internal gauge. Per-example contribution MSE is computed on centered contributions over the validation split. See… See the full description on the dataset page: https://huggingface.co/datasets/arushram/fnbm-current-gt-motif-effects-sensitivity-20260729.motif_library_final
Motif Modules V9 (Black Line Version)
Each row contains a motif pair:
png: raster preview (512×512)
svg: vector paths of the same motif
Total: 100 pairsCreated: 2025-10-06Author: maryzhangLicense: CC-BY-NC-SA 4.0
Chinese Porcelain Motif Library - Final Collection (PNG & SVG)
Dataset Description
Dataset Summary
A curated collection of 100 high-quality Chinese porcelain motifs provided in both raster (PNG) and vector (SVG) formats. These motifs have… See the full description on the dataset page: https://huggingface.co/datasets/maryzhang/motif_library_final.
