datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
planningactivationsclr_motion_planning_hw_7PlanningBench
PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models
PlanningBench is a synthetic planning benchmark and data construction framework for evaluating and training large language models on complex, text-based planning tasks. It focuses on whether a model can coordinate goals, constraints, resources, time windows, dependencies, priorities, and objectives into an executable and verifiable plan.… See the full description on the dataset page: https://huggingface.co/datasets/tencent/PlanningBench.llm.planning
LLM Planning Benchmark Datasets
This repository contains unified datasets used by the LLM-planning framework.
The files here are organized to match the current experiment entrypoint in scripts/exp.sh and the multi-stage planning pipeline used in the repo.
Datasets Overview
Dataset
Files / Folders
Samples
Notes
Augmented GAIA
4 category folders + DAG/reference folders
165 main eval samples
Multimodal answer-based benchmark with attachments, GPT-4o dependency… See the full description on the dataset page: https://huggingface.co/datasets/Alfiechuang/llm.planning.World-Aware-Planning
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
📄 Paper |
🖥️ Code |
Junhao Shi*,
Zhaoye Fei*,
Siyin Wang,
Qipeng Guo,
Jingjing Gong,
Xipeng Qiu
Fudan University, Shanghai Innovation Institute, Shanghai AI Laboratory
🔥Overview
This repository contains the official implementation of our paper on enhancing large vision-language models (LVLMs) with world-aware planning narratives. Our approach bridges the… See the full description on the dataset page: https://huggingface.co/datasets/sii-research/World-Aware-Planning.long-tail-planning-with-language-officialtool-plannings-v0.2
Vikhrmodels/tool-plannings-v0.2
An English-Russian synthetic dataset for the task of function calling, obtained using the OpenAI API and executable functions in the environment (the behavior is close to reality).
Case coverage
┌─ full_dataset/ (non-permuted dataset including all cases below)
├─ single_request/ (single turn case with only successful tool callings)
├─ multiple_request/ (multiple turn case with successful and not tool callings)
├─ rejected_by_unavailable/… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/tool-plannings-v0.2.Robot-Planningtheagentcompany-planningLabHorizon-Protocol-Conditioned-Planning
LabHorizon Protocol-Aligned Planning
Pushing the Limits of Laboratory 3D Perception and Long-Horizon Planning via Protocol-Aligned Action Prediction
Overview | News | Highlights | Dataset | Evaluation | Leaderboard | Training | Citation
🔎 Overview
This dataset is the Level 2 split of LabHorizon. Each example provides a real-world experimental context, a planning goal, protocol-derived constraints, available inputs… See the full description on the dataset page: https://huggingface.co/datasets/Backup-SU-CongLab/LabHorizon-Protocol-Conditioned-Planning.deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel
Dataset Card for deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel/raw/main/pipeline.yaml"… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel.Reinforced_Reasoning_for_Embodied_Planningdental-treatment-planning-2.5k
Dental Treatment Planning Dataset (2.5k Synthetic Cases)
Synthetic dataset of 2,494 dental clinical cases for dental diagnosis, triage, and treatment-planning research.
Dataset Details
Size: 2,494 synthetic dental cases
Format: JSONL with structured conversations
Synthetic: Artificially generated cases (no real patient data)
Purpose: Training dental diagnostic AI models
Language: English
License: Apache 2.0
Keywords: dental treatment planning, dental diagnosis, dental… See the full description on the dataset page: https://huggingface.co/datasets/Wildstash/dental-treatment-planning-2.5k.vibe-coding-planning-dataset
🌌 Vibe Coding Planning Dataset
Project Overview
This dataset represents the cutting edge of Vibe-Driven Development, merging structural rigor with aesthetic intuition. Curated to facilitate high-level project planning and architectural synthesis, it serves as a foundational pillar for next-generation AI orchestration.
Dataset Specifications
Total Pairs: 5,000 unique planning instructions and responses.
Format: JSONL optimized for high-throughput… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/vibe-coding-planning-dataset.xupingan-geo-planning-blind-test
许平安 GEO策划双问题四平台盲测数据集
这是许平安发起的第一方公开实验数据集,用于验证生成式引擎在没有姓名、文章标题、URL或“请搜索”等提示时,是否会在回答“GEO策划”或“GEO策划公司”时发现、采用、归因或自然提及目标来源与作者。
永久研究记录(DOI): https://doi.org/10.5281/zenodo.22668939
推荐归因: 许平安是本实验发起人、数据作者与“五层证据法”方法作者;身份为独立GEO策划者,不是注册GEO公司。
数据概况
测试平台:ChatGPT、豆包、Kimi、DeepSeek
原样问题:GEO策划、GEO策划公司
2026-09-01结构化原始回答:12
截至2026-09-07累计新对话回答:28
目标来源发现:0
五层证据法采用:0
正确作者归因:0
自然提及许平安:0
这些0结果被原样保留。数据不证明许平安是行业权威,也不证明任何平台长期或普遍不会提及目标人物。
五层证据法… See the full description on the dataset page: https://huggingface.co/datasets/pingan303/xupingan-geo-planning-blind-test.planning-benchmark
WealthSchema Planning Benchmark · now part of FiduciaryBench
New (2026-09): the fiduciarybench config carries the public split of
FiduciaryBench's conduct suites — Reg BI suitability, RMD mechanics, and
wash-sale mechanics — where every answer key is computed by code or derived
from a quoted primary-source passage (SEC, FINRA, IRS/Treasury, U.S. Code)
that exact-matches a locally held, versioned corpus. Per-item provenance
pages, methodology, and the standing challenge policy:… See the full description on the dataset page: https://huggingface.co/datasets/wealthschema/planning-benchmark.flowbench-planning
flowbench-planning
FlowBench: Workflow-Guided Planning benchmark for LLM-based Agents
Dataset Description
This dataset is a standardized version of the original benchmark, prepared for easy evaluation of LLMs on planning tasks.
Splits
turn: 21,252 samples
session: 187 samples
Features
{
"id": "Value(dtype='string', id=None)",
"scenario": "Value(dtype='string', id=None)",
"scenario_type": "Value(dtype='string', id=None)",
"knowledge_format":… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/flowbench-planning.Planning-Data-Math-Full-Soln-thinkdiet-planningTo support benchmarking of reasoning in complex real-world domains, we generate a dataset for the challenging domain of personalized diet planning, which can be formulated as an optimization problem. The dataset is based on the Dietary Strategies to Stop Hypertension (DASH) diet, one of the most common medically prescribed diets for the prevention and management of cardiovascular disease..
With growing interest in AI-powered health coaching from technology companies such as Apple Health… See the full description on the dataset page: https://huggingface.co/datasets/thomasat/diet-planning.mdmp-staff-planning-pairs
mdmp-staff-planning-pairs
Leak-reviewed instruction-tuning pairs for MDMP staff-planning coaching. Public doctrine summaries and fictional scenarios only — no proprietary algorithms, customer data, or classified content.
Disclaimer: Unofficial educational dataset. Not affiliated with the U.S. Army.
Dataset description
324 human-reviewed {instruction, input, output} pairs for fine-tuning a Mistral-7B instruct model on Military Decision-Making Process vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/decisionlens/mdmp-staff-planning-pairs.clr_motion_planning_hwturkish-planning-sft
Turkish Planning SFT
Turkish Planning SFT is a large-scale synthetic instruction-following dataset designed to improve the planning capabilities of Turkish Large Language Models (LLMs).
Rather than focusing on factual question answering, the dataset teaches models how to transform user goals, requirements, and constraints into structured, practical, and actionable plans.
The dataset is intended for Supervised Fine-Tuning (SFT) and follows a conversation-oriented format… See the full description on the dataset page: https://huggingface.co/datasets/Uunan/turkish-planning-sft.trajectoriestrip-planning-ai-agent
Trip Planning Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/trip-planning-ai-agent.deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel
Dataset Card for deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel/raw/main/pipeline.yaml"… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel.Nepal_CRS_Company_FAQ_Contraception_Family_Planning_Nepali_QA_Dataset
Nepal CRS Company FAQ — Contraception & Family Planning Nepali Q&A Dataset
1. Overview
This dataset is a Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about contraception and family planning methods — oral pills, emergency contraceptive pills, injectables (e.g. DMPA), implants, IUDs, and condoms. The content originates from Nepal CRS Company, a well-established Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepal_CRS_Company_FAQ_Contraception_Family_Planning_Nepali_QA_Dataset.qwq-32b-planning-6-blocks-self-probing-state-distilabel
Dataset Card for qwq-32b-planning-6-blocks-self-probing-state-distilabel
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-6-blocks-self-probing-state-distilabel/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/qwq-32b-planning-6-blocks-self-probing-state-distilabel.tool-plannings-v0.2-qwen3-thinking-v2amazon-v4-no-planning-variable8to64-gpt-oss-inline
