datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
motion-smd-data
Motion-SMD Data
Data release for "Encoder-Free Human Motion Understanding via Structured Motion Descriptions".
🌐 Project page: https://yaozhang182.github.io/motion-smd/
💻 Code: https://github.com/yaozhang182/motion-smd
🤗 LoRA adapters: https://huggingface.co/zyyy12138/motion-smd-lora
📄 Paper (arXiv): https://arxiv.org/abs/2604.21668
What's here
Four subdirectories, each with its own README.md describing files, provenance, and license:
Subdir
Contents
Our… See the full description on the dataset page: https://huggingface.co/datasets/zyyy12138/motion-smd-data.MotionSightThis is the dataset proposed in our paper MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs.
We split the dataset into multiple small files, you can recover by cat:
cat MotionSightDataset_part* > MotionSightDataset.zip
unzip MotionSightDataset
Project Page | Github
motivational_quotes
Motivational Quotes for Reservists
This dataset contains 1,000+ AI-generated motivational quotes, each categorized by theme such as resilience, courage, discipline, and perseverance.It was created to support Israeli reserve soldiers (“Miluim”) by offering uplifting, emotionally impactful messages during active service and difficult times.
🧾 Dataset Details
Created by: AMaACHINE
Language(s): English
License: OpenRAIL
Model Used: google/flan-t5-base from Hugging… See the full description on the dataset page: https://huggingface.co/datasets/AMaACHINE/motivational_quotes.task294_storycommonsense_motiv_text_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task294_storycommonsense_motiv_text_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task294_storycommonsense_motiv_text_generation.alpaca_ccass_motivations_sommaires_titres
Training dataset for summarizing and titling decisions of the French Court of cassation based on motivations
This alpaca-format dataset is designed to train models for summarizing and titling French Supreme Court decisions based on the grounds of them. Created with a view to producing metadata for decisions not published in the bulletin, this dataset aims to simplify the development of annotation and categorization tools, and is positioned as a facilitator for jurisprudential… See the full description on the dataset page: https://huggingface.co/datasets/Cour-de-cassation/alpaca_ccass_motivations_sommaires_titres.EarthScience-Text-LLM-20K-90-10
EarthScience-Text-LLM-20K-90-10
This is a pure-text Earth-science corpus unified from three non-overlapping upstream datasets:
Ekimetrics/climateqa-ipcc-ipbes-reports-1.0: climate and IPCC/IPBES report chunks.
GeoGPT-Research-Project/GeoGPT-CoT-QA: geoscience question-answer reasoning.
gremlin97/RemoteSensingCorpus: remote-sensing and geospatial machine-learning text.
Files and Split
The previous preprocessing outputs were merged into a 23,098-record pool and… See the full description on the dataset page: https://huggingface.co/datasets/moTcream/EarthScience-Text-LLM-20K-90-10.motivational-quotes
Dataset Card for Motivational Quotes
This is a dataset of motivational quotes, scraped from Goodreads. It contains more than 4000 quotes, each of them labeled with the corresponding author.
Data overview
The quotes subset contains the raw quotes and the corresponding authors. The quotes_extended subset contains the raw quotes plus a short prompt that can be used to train LLMs to generate new quotes:
// quotes
{
"quote": "“Do not fear failure but rather fear not… See the full description on the dataset page: https://huggingface.co/datasets/asuender/motivational-quotes.AI-Agent-Generating-Tool-Debugging-Prompt-Library
Dataset Card for "AI Agent Generating Tool & Debugging Prompt Library" 🤖⚙️
Dataset Details 📚
Dataset Name: AI Agent Generating Tool & Debugging Prompt Library
Dataset Description:This dataset includes a collection of prompts focused on building and debugging AI-driven tools, including creating self-improving AI agents and debugging prompts for Python projects. The dataset is designed for use in fine-tuning models related to code generation, debugging, and software… See the full description on the dataset page: https://huggingface.co/datasets/Chemically-motivated/AI-Agent-Generating-Tool-Debugging-Prompt-Library.motivational-quotes
Mentria Motivational Quotes
581 hand-curated, original motivational quotes, written and curated as LoRA
fine-tuning data for the quote generator at
mentria.ai/tools/quote. Every line was either
written by hand for this dataset or individually reviewed before inclusion —
no scraped content, no famous quotes in disguise.
Diversity engineering
Style-skewed training data drags LoRA adapters into a single template, so this
set was built with enforced diversity quotas… See the full description on the dataset page: https://huggingface.co/datasets/mentriaai/motivational-quotes.Gitruck-MotionIR
Gitruck MotionIR
Gitruck MotionIR is a Chinese motion-design dataset that aligns project-level
natural-language descriptions, technique-level annotations, temporal evidence,
and a renderable intermediate representation (IR v1). The corpus was normalized
from authorized Alight Motion, After Effects, NodeVideo, and Jianying projects.
Gitruck MotionIR 是一个中文动效设计数据集,将工程级描述、技法级标注、时间证据与可渲染
IR v1 对齐。语料由已获授权的 Alight Motion、After Effects、NodeVideo 与剪映工程归一化而来。
Dataset summary /… See the full description on the dataset page: https://huggingface.co/datasets/Hocassian/Gitruck-MotionIR.MoT-Code-350K
🏠 MoTCode-Data
• 🤗 Data • 🤗 Model • 🐱 Code • 📃 Paper
Dataset Structure
from datasets import load_dataset
load_dataset("JingyaoLi/MoT-Code-350K")
DatasetDict({
train: Dataset({
features: ['instruction', 'output'],
num_rows: 312645
})
})
Modular-of-thought Data Creation
We provide an example python file to evolution a MoT dataset. Run the following command:
python src/generate_MoT_dataset.py \
--data_path $data_path \… See the full description on the dataset page: https://huggingface.co/datasets/JingyaoLi/MoT-Code-350K.fineweb-ultra-mini-pro
Dataset Card for Fineweb Ultra Mini
Fineweb Ultra Mini is a dataset derived from the original Fineweb dataset made by huggingface (see here: https://huggingface.co/datasets/HuggingFaceFW/fineweb).
The dataset focuses on extracting high quality data from the Fineweb dataset, from the 1-0.5% range. If you would like more data, though slightly sacrificing quality check out fineweb ultra mini, which focuses on the 2-3% of high quality data originally found in fineweb.… See the full description on the dataset page: https://huggingface.co/datasets/motionlabs/fineweb-ultra-mini-pro.Motivation_Employee_Engagement_Content_1
Motivation Employee Engagement Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Motivation_Employee_Engagement_Content_1.Motivacion-Diaria
Dataset Card for Dataset Name
Name
Motivación Diaria
Dataset Summary
Scrapeado de http://www.motivaciondiaria.com/
Languages
[Spanish]
full-html-stying-dataset-motion-geometry-styles
Full HTML Tailwind Motion Geometry Styles
kogai/full-html-stying-dataset-motion-geometry-styles contains phase1_dataset_10k_motion_geometry_style.jsonl, a JSONL dataset with 10000 synthetic examples. Tailwind styling transforms emphasizing motion, geometry, and expressive visual direction for full HTML pages.
Schema
input_html: source HTML before Tailwind classes are added.
output_html: transformed HTML with Tailwind utility classes applied.
style_phrase:… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-motion-geometry-styles.Motivation_Employee_Engagement_Content_2
Motivation Employee Engagement Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Motivation_Employee_Engagement_Content_2.wisconsin-motorists-handbook-dataset
Wisconsin Motorists Handbook Dataset
Generated by DocParserEngine.
Field
Value
Documents
1
Records
1
Schema
full
Usage
from datasets import load_dataset
ds = load_dataset("Remixonwin/wisconsin-motorists-handbook-dataset")
Motor_LLM_datasetashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak
Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts, drop مجزوء الوافر, drop المتدارك)
This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer and keeps the same schema, columns, locked prompt format, and general structure as the upstream phase-1 dataset.
The only additional change is the removal of rows where:
base_meter == "المتدارك"
This removes all poems whose base meter is المتدارك from the published… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak.everyday-user-queriesdeepseek-r1-distill-llama-70b-synthetic
Dataset Card for Dataset Name
This dataset contains input and output pairs from the popular Deepseek-R1 model, distilled into Llama-70B.
Dataset Details
Curated by: ReflexAI
Language(s) (NLP): English
License: Llama3.3
If you like the open source work of ReflexAI, don't hesitate to give us a follow on huggingface, or like this dataset.
