datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Motion-o-MCoT-PLM-motion-keyframes
Motion-o-MCoT (PLM + motion keyframes)
Subset of STGR: STR_plm_rdcap rows with <motion in reasoning_process, plus sharded keyframes under videos/stgr/plm/kfs/.
Train split: 3,168 examples (see export_manifest.json in the repo for exact export stats).
Keyframes: JPEGs are stored under shard subfolders (e.g. videos/stgr/plm/kfs/plm_0150/…) so each directory stays under Hugging Face file-count limits. Each key_frames[].path in the JSON is relative to videos/stgr/plm/kfs/ (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bishoygaloaa/Motion-o-MCoT-PLM-motion-keyframes.StarCraft-Motionmotif-qa
MotifQA
Dataset Summary
MotifQA is a synthetic graph question-answering benchmark focused on detecting graph motifs inside small random graphs.
Each example pairs a textual prompt with an answer sentence, a list of nodes highlighted as the motif (when present), and an explicit graph description(nodes and edges).
In this QA dataset, all graphs are homogenous and undirected.
Subsets cover both yes/no motif detection, motif-type classification (house vs 5-cycle), and… See the full description on the dataset page: https://huggingface.co/datasets/naos-ku/motif-qa.MotionFix
MotionFix MotionHub Format
This dataset contains the processed MotionFix motion-editing data in the MotionHub format.
It stores paired source and target motions as SMPL-H 52-joint parameter files, plus editing instructions and split annotations.
Please also follow the license and terms of the original MotionFix dataset.
Structure
MotionFix/
├── smplh_52/
│ ├── train/
│ │ ├── 000000_000499/
│ │ ├── 000500_000999/
│ │ └── ...
│ ├── val/
│ └── test/… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/MotionFix.SIS-Motion-54K
✨SIS-Motion-54K✨
SIS-Motion-54K is a motion-aware instruction-tuning dataset built from the AirScape training split. It is designed to fine-tune the SIS-Motion model for joint understanding of space (environment) and self (agent motion) in embodied UAV scenarios.
Important: This dataset is strictly separated from the SIS-Bench evaluation benchmark. It contains only perception and memory tasks — no reasoning-level data — so any generalization gains on… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/SIS-Motion-54K.github-issuesCyberSecurityDataset
Dataset Card for Cyber Security Dataset
This dataset provides a collection of curated data points related to cybersecurity, focusing on penetration testing, known exploits, and vulnerability analysis. It is intended to aid researchers, educators, and developers in building AI tools for cybersecurity applications.
Dataset Details
Dataset Description
This dataset contains labeled information about exploits, vulnerabilities, and penetration testing techniques.… See the full description on the dataset page: https://huggingface.co/datasets/Chemically-motivated/CyberSecurityDataset.duckdb-text2sql-25k
Dataset Summary
The duckdb-text2sql-25k dataset contains 25,000 DuckDB text-2-sql pairs covering diverse aspects of DuckDB's SQL syntax.
We synthesized this dataset using Mixtral 8x7B, based on DuckDB's v0.9.2 documentation and Spider schemas that were translated to DuckDB syntax and enriched with nested type columns.
Each training sample consists of a natural language prompt, a corresponding (optional) schema, and a resulting query. Each pair furthermore has a category property… See the full description on the dataset page: https://huggingface.co/datasets/motherduckdb/duckdb-text2sql-25k.alpaca_ccass_motivations_sommaires_titres
Training dataset for summarizing and titling decisions of the French Court of cassation based on motivations
This alpaca-format dataset is designed to train models for summarizing and titling French Supreme Court decisions based on the grounds of them. Created with a view to producing metadata for decisions not published in the bulletin, this dataset aims to simplify the development of annotation and categorization tools, and is positioned as a facilitator for jurisprudential… See the full description on the dataset page: https://huggingface.co/datasets/Cour-de-cassation/alpaca_ccass_motivations_sommaires_titres.motor_temperature_monitoring.jsonmotionatlas-bench
MotionAtlas-Bench v1
MotionAtlas-Bench v1 is a video multiple-choice benchmark for motion and target-entity understanding. This public release contains MCQ records, answer keys, media, and target-object masks needed to reproduce the visual grounding settings.
Resources
Paper: https://arxiv.org/abs/2606.29531
Project page: https://kagura-0001.github.io/projects/MotionAtlas/
Code: https://github.com/Kagura-0001/MotionAtlas
MotionAtlas-Data:… See the full description on the dataset page: https://huggingface.co/datasets/maxLWSv2/motionatlas-bench.EarthScience-Text-LLM-20K-90-10
EarthScience-Text-LLM-20K-90-10
This is a pure-text Earth-science corpus unified from three non-overlapping upstream datasets:
Ekimetrics/climateqa-ipcc-ipbes-reports-1.0: climate and IPCC/IPBES report chunks.
GeoGPT-Research-Project/GeoGPT-CoT-QA: geoscience question-answer reasoning.
gremlin97/RemoteSensingCorpus: remote-sensing and geospatial machine-learning text.
Files and Split
The previous preprocessing outputs were merged into a 23,098-record pool and… See the full description on the dataset page: https://huggingface.co/datasets/moTcream/EarthScience-Text-LLM-20K-90-10.maniskill-pickcube-motionplanning-lerobot-v3
ManiSkill PickCube motion-planning — verified LeRobot v3 reference subset
Need to turn your own raw robotics data into LeRobot v3.0? Start a free conversion →
Community conversion produced by ViaCatalyst BYOD. This repository is not an official upstream release and is not affiliated with the ManiSkill authors.
This is a compact, provenance-complete conversion of the first 10 trajectories in the pinned ManiSkill PickCube-v1 motion-planning HDF5 file. It is a reproducible… See the full description on the dataset page: https://huggingface.co/datasets/ViaCatalyst/maniskill-pickcube-motionplanning-lerobot-v3.LLaVA-Pretrain-JA
Dataset Details
Dataset Type:Japanese LLaVA Pretrain is a localized version of the original LLaVA Pretrain dataset. This version is translated into Japanese using DeepL API and is aimed at serving similar purposes in the context of Japanese language.
Resources for More Information:For information on the original dataset: LLaVA
License:License: Must comply with license of CC-3M, BLIP (if you use their synthetic caption).
CC-3M The dataset may be freely used for any purpose, although… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/LLaVA-Pretrain-JA.motivational-quotes
Dataset Card for Motivational Quotes
This is a dataset of motivational quotes, scraped from Goodreads. It contains more than 4000 quotes, each of them labeled with the corresponding author.
Data overview
The quotes subset contains the raw quotes and the corresponding authors. The quotes_extended subset contains the raw quotes plus a short prompt that can be used to train LLMs to generate new quotes:
// quotes
{
"quote": "“Do not fear failure but rather fear not… See the full description on the dataset page: https://huggingface.co/datasets/asuender/motivational-quotes.motibench
MotiBench
MotiBench is a benchmark for evaluating video generation models under physically grounded and commonsense-driven settings. Each image depicts a moment immediately before a physical event, in which a small, localized action is expected to trigger a larger physical response. All images explicitly capture the pre-event state, in which no visible motion has yet occurred, yet the physical configuration strongly implies an imminent interaction.
Sources and Task… See the full description on the dataset page: https://huggingface.co/datasets/shinying/motibench.motivational-quotes
Mentria Motivational Quotes
581 hand-curated, original motivational quotes, written and curated as LoRA
fine-tuning data for the quote generator at
mentria.ai/tools/quote. Every line was either
written by hand for this dataset or individually reviewed before inclusion —
no scraped content, no famous quotes in disguise.
Diversity engineering
Style-skewed training data drags LoRA adapters into a single template, so this
set was built with enforced diversity quotas… See the full description on the dataset page: https://huggingface.co/datasets/mentriaai/motivational-quotes.MotiveBench
MotiveBench
This is the official repository for our ACL 2025 paper "MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?"
Dataset Description
MotiveBench is a benchmark for evaluating the human-like motivational and behavioral reasoning capabilities of large language models (LLMs). It consists of 200 diverse profiles and 600 reasoning tasks, covering multiple levels of motivation based on Maslow's Hierarchy of Needs. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/chicosirius/MotiveBench.motor-fuel-dispenser-measurement-tolerances-nist
How far off a US fuel pump, LPG meter or EV charger is allowed to be (NIST Handbook 44)
Canonical, always-current version: https://referencesource.org/motor-fuel-dispenser-measurement-tolerances-nist/
Machine-readable: https://referencesource.org/motor-fuel-dispenser-measurement-tolerances-nist/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-20
Stale after: 2027-08-20 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/motor-fuel-dispenser-measurement-tolerances-nist.reachy-mini-motion-synth
Reachy Mini Text-to-Motion (real + synthetic)
Expressive motions for Reachy Mini: each clip pairs a
text prompt (an emotion, reaction, character or situation) with a head / antenna / body-yaw trajectory.
Built to train and evaluate text → motion models that generalise beyond the handful of real recordings.
source
clips
duration
real: Pollen emotions library, reachability-fixed
85
8.6 min
generated_batch1: 154 prompts × 4 plan variants × 2 render seeds
1,232
98 min… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/reachy-mini-motion-synth.motor-fuel-minimum-markup-laws-by-state
State below-cost motor fuel selling laws and minimum-markup percentages
Canonical, always-current version: https://referencesource.org/motor-fuel-minimum-markup-laws-by-state/
Machine-readable: https://referencesource.org/motor-fuel-minimum-markup-laws-by-state/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-20
Stale after: 2027-02-16 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 18
A handful… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/motor-fuel-minimum-markup-laws-by-state.outboard-motor-service-specs
Outboard motor service specifications by brand and model
Canonical, always-current version: https://referencesource.org/outboard-motor-service-specs/
Machine-readable: https://referencesource.org/outboard-motor-service-specs/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-05
Stale after: 2028-08-04 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 53
Spark plug, gap, oil capacity, oil filter, gear… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/outboard-motor-service-specs.duckdb-docbench
DocBench: A Synthetic DuckDB Text-to-SQL Benchmark
DocBench is a synthetic Text-to-SQL benchmark dataset consisting of 2430 question/sql pairs derived from the DuckDB documentation, specifically designed to probe language models for knowledge of DuckDB-specific SQL functionality.
The dataset covers functions, aggregates, operators, statements, keywords, and multi-keyword expressions available in DuckDB 1.1.3 and its default extensions.
Dataset Structure
Each example… See the full description on the dataset page: https://huggingface.co/datasets/motherduckdb/duckdb-docbench.cdg-motorcycle-repair-qa-data-85x10
Dataset Card for Motorcycle Repair QA Dataset (Generated)
Dataset Summary
This dataset contains 880 question-answer pairs focused on various aspects of motorcycle repair and maintenance. The data was synthetically generated using the meta-llama/Meta-Llama-3-8B-Instruct model via the script. The generation process aimed to create relevant and informative QA pairs based on a list of 85 specific topics related to motorcycle repair.
The dataset is provided in JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-motorcycle-repair-qa-data-85x10.pg-docbench
DocBench: A Synthetic PostgreSQL Text-to-SQL Benchmark
DocBench is a synthetic Text-to-SQL benchmark dataset consisting of 4398 question/sql pairs derived from the PostgreSQL documentation, specifically designed to probe language models for knowledge of PostgreSQL-specific SQL functionality.
The dataset covers functions, aggregates, operators, statements, keywords, and multi-keyword expressions available in PostgreSQL 14 and PostgreSQL 18, including extensions like pgvector.… See the full description on the dataset page: https://huggingface.co/datasets/motherduckdb/pg-docbench.MotionChain_Conv
MotionChain: Conversational Motion Controllers via Multimodal Prompts
MotionChain introduces a multi-modal human motion conversation dataset with support for multi-modal prompts across diverse motion tasks.
Data Preparation
The whole MotionChain dataset comprises two main components: Human Motion data and language annotations.
Step 1. Download and Prepare the Human Motion Data.
Prepare human motion data from HumanML3D.
Follow the instructions HumanML3D and download the… See the full description on the dataset page: https://huggingface.co/datasets/OpenMotionLab/MotionChain_Conv.han-embodied-motion-intention-dataset-v1
Embodied Motion Intention Dataset
This dataset maps intended actions to physical motion
for humanoid agents operating in real or simulated environments.
Use Cases
Motion planning
Intent-to-action mapping
Safer physical interaction
Fields
intention_label
environment_state
motion_priority
safety_flag
Part of
Humanoid Network (HAN)
License
MIT
lead_marketing_runtime_Jan_2026motivational-english-quotes
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/aldoyh/motivational-english-quotes.motion_planning_for_automated_drivingThis question-answer dataset is extracted from paper A Survey on Hybrid Motion Planning Methods for Automated Driving Systems written by MReza Alipour Sormoli, Konstantinos Koufos, Mehrdad Dianati, Senior Member, IEEE, and Roger Woodman. Deepseek R1 is used as the extraction tool.
