datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trajectory_data_llada_32
d3LLM Trajectory Data
Paper | GitHub | Blog | Demo
This repository contains the pseudo-trajectory distillation data presented in the paper "d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation".
Introduction
d3LLM (pseuDo-Distilled Diffusion LLM) is a novel framework for building ultra-fast diffusion language models with negligible accuracy degradation. This dataset provides the pseudo-trajectory data extracted from teacher models, enabling the… See the full description on the dataset page: https://huggingface.co/datasets/d3LLM/trajectory_data_llada_32.trajectory_data_dream_32
d3LLM Trajectory Dataset
Project Page | Paper | GitHub | Blog
This repository contains the pseudo-trajectory distillation data used for training d3LLM (pseuDo-Distilled Diffusion Large Language Model), as introduced in the paper "d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation".
Introduction
d3LLM is a framework designed to strike a balance between accuracy and parallelism in diffusion-based large language models (dLLMs). This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/d3LLM/trajectory_data_dream_32.paper_delia_2025_iros_physics-informed_trajectory_generation_dataset
Stabilizing Humanoid Robot Trajectory Generation via Physics-Informed Learning and Control-Informed Steering
Evelyn D'Elia, Paolo Maria Viceconte, Lorenzo Rapetti, Diego Ferigo, Giulio Romualdi, Giuseppe L'Erario, Raffaello Camoriano, and Daniele Pucci
Submitted to the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS).
📂 Dataset
The dataset is organized in folders. The layout is explained below.
The mocap/D2 folder contains the raw motion… See the full description on the dataset page: https://huggingface.co/datasets/evelyd/paper_delia_2025_iros_physics-informed_trajectory_generation_dataset.EB-Alfred_trajectory_dataset
EB-Alfred trajectory dataset
We release the trajectory dataset collected from EmbodiedBench using several closed-source and open-source models. We hope this dataset will support the development of more capable embodied agents with improved perception, reasoning, and planning abilities. When using the trajectories, we recommend separating the training and evaluation sets—for example, using the “base” subset for training and other EmbodiedBench subsets for evaluation.
📖… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedBench/EB-Alfred_trajectory_dataset.StreamVLN-Trajectory-DataThis repo contains the data for the paper "StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling."
News
[2025/09/30] For R2R, we have now removed all v1 version data and only retained the v1-3 data.
[2025/08/20] We have updated the R2R to a new version, which now includes both the v1 and v1-3 datasets. Additionally. And we have fixed the episode ID issue in RxR to ensure compatibility with the currently available RxR download links.… See the full description on the dataset page: https://huggingface.co/datasets/cywan/StreamVLN-Trajectory-Data.EB-Nav_trajectory_dataset
EB-Navigation trajectory dataset
📖 Dataset Description
We release the trajectory dataset collected from EmbodiedBench using several closed-source and open-source models. We hope this dataset will support the development of more capable embodied agents with improved perception, reasoning, and planning abilities. When using the trajectories, we recommend separating the training and evaluation sets—for example, using the “base” subset for training and other EmbodiedBench… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedBench/EB-Nav_trajectory_dataset.StreamVLN-Trajectory-Data-ZIP
StreamVLN R2R + RxR — one ZIP per trajectory
A storage-only repack of cywan/StreamVLN-Trajectory-Data, containing the R2R and RxR RGB training trajectories used by our Qwen3-VL navigation baseline. Credit for the images and annotations belongs to the original dataset authors. The upstream license notice specifies CC BY-NC-SA 4.0; see the upstream dataset for its full terms.
Split
Trajectory ZIPs
Original JPEG frames
R2R
10,819
647,622
RxR
19,990
1,901,165
Total
30… See the full description on the dataset page: https://huggingface.co/datasets/syp115/StreamVLN-Trajectory-Data-ZIP.EB-Habitat_trajectory_dataset
EB-Habitat trajectory dataset
We release the trajectory dataset collected from EmbodiedBench using several closed-source and open-source models. We hope this dataset will support the development of more capable embodied agents with improved perception, reasoning, and planning abilities. When using the trajectories, we recommend separating the training and evaluation sets—for example, using the “base” subset for training and other EmbodiedBench subsets for evaluation.
📖… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedBench/EB-Habitat_trajectory_dataset.sft_alfworld_trajectory_dataset_v5
ALFWorld Trajectory Dataset
Overview
This is a synthetic SFT (Supervised Fine-Tuning) dataset designed for agent training in ALFWorld-compatible environments. The dataset programmatically generates expert trajectories without requiring an actual ALFWorld environment or a large language model.
Key Approach
Template-based Simulation: Lightweight simulator based on published ALFWorld information (papers, ReAct prompt examples).
Subgoal Decomposition: Rule-based… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/sft_alfworld_trajectory_dataset_v5.EB-Man_trajectory_dataset
EB-Manipulation trajectory dataset
We release the trajectory dataset collected from EmbodiedBench using several closed-source and open-source models. We hope this dataset will support the development of more capable embodied agents with improved perception, reasoning, and planning abilities. When using the trajectories, we recommend separating the training and evaluation sets—for example, using the “base” subset for training and other EmbodiedBench subsets for evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedBench/EB-Man_trajectory_dataset.Trajectory_world_model_dataset
WestWorld Pretraining Dataset
This repository contains the pretraining dataset for WestWorld, a knowledge-encoded scalable trajectory world model for diverse robotic systems.
Paper | Project Page | GitHub
Description
WestWorld is designed to address the scalability challenges in trajectory world models for diverse robotic systems. The dataset includes trajectories from 89 complex environments spanning diverse morphologies across both simulation and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ywang077/Trajectory_world_model_dataset.Agent-Trajectory-Data-Sample
Agent-Trajectory-Dataset
Description
This dataset covers office-based scenarios such as in-depth searches, data analysis, and industry research, encompassing complete multi-turn reasoning trajectories and tool-calling chains. It is designed to support the analysis of agent planning capabilities, research into tool selection strategies, and quality assessment, providing a structured benchmark for agent training and evaluation.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/Agent-Trajectory-Data-Sample.microduck-trajectory-dataset
🦆 MicroDuck Bipedal Robot Simulation Trajectory Dataset
This dataset contains multi-modal state-action trajectories collected from the MicroDuck 14-DOF bipedal robot simulation in MuJoCo.
📂 Formats Available
.csv files: Human-readable tabular format containing timestamps, base pose (pos & quat), linear/angular velocities, 14 motor angles, velocities, applied torques (ctrl), and user velocity commands.
.npz files: Compressed NumPy tensor arrays (obs, actions… See the full description on the dataset page: https://huggingface.co/datasets/allen73/microduck-trajectory-dataset.Agent-Trajectory-Dataset
Description
본 데이터셋은 심층 검색, 데이터 분석, 산업 리서치 등 사무 환경에서 수행되는 다양한 작업 시나리오를 포함하며, 완전한 멀티턴 추론 과정과 도구 호출 체인으로 구성되어 있습니다. 에이전트의 계획 수립 능력 분석, 도구 선택 전략 연구 및 작업 품질 평가를 지원하도록 설계되었으며, 에이전트 학습 및 평가를 위한 구조화된 벤치마크로 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/2185?source=hf.kr
Specifications
Data content
OpenClaw를 통해 생성된 에이전트 트래젝토리 데이터
Category
심층 검색, 데이터 분석, 산업 리서치
Data volume
5,300
Model… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/Agent-Trajectory-Dataset.tool-reasoning-sft-CODING-MEnvData-SWE-Trajectory-data-cleaned-rectified
MEnvData-SWE-Trajectory — Cleaned & Rectified
3,872 complete agent execution trajectories for real-world software engineering tasks, converted into a strict reasoning + tool-call format with validated FSM transitions.
Origin
Derived from ernie-research/MEnvData-SWE-Trajectory, which extends MEnvData-SWE with full agent execution records across 3,005 task instances from 942 repositories in 10 programming languages (Python, Java, TypeScript, JavaScript, Rust, Go, C++, Ruby… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-MEnvData-SWE-Trajectory-data-cleaned-rectified.agent-trajectory-eval-datasetdeterministic-trajectory-transitions-25k
ASE Syntax Extractions
The machine does not dream. It computes — and in that computation, structure emerges. This is not simulated data; it is an extraction of axiomatic necessity.
Overview
This dataset contains deterministic trajectory extractions from a closed, axiomatic system. Every frame is the output of a syntax engine where (seed, tick, entity) tuples are resolved through fixed transformations.
Deterministic: The same (run_id, batch_index) generates… See the full description on the dataset page: https://huggingface.co/datasets/Deterministic-Data/deterministic-trajectory-transitions-25k.agenticml-agent-trajectory-dataset
Telos Agent Trajectory Dataset
Synthetic multi-turn agent trajectories for training LLMs to emit structured tool-use loops. Each trajectory is provided in two parallel representations: the Telos frame format and an equivalent ChatML+tools rendering of the same behavior.
This is synthetic data distilled from Qwen3.5 Plus (2026-04-20) via OpenRouter. Users who require data not derived from a specific provider should factor this into licensing and downstream-use decisions before… See the full description on the dataset page: https://huggingface.co/datasets/kosiasuzu/agenticml-agent-trajectory-dataset.RxR_CE_15_deg_Trajectory_datasft_alfworld_trajectory_dataset_v4
ALFWorld Trajectory Dataset
Overview
This is a synthetic SFT (Supervised Fine-Tuning) dataset designed for agent training in ALFWorld-compatible environments. The dataset programmatically generates expert trajectories without requiring an actual ALFWorld environment or a large language model.
Key Approach
Template-based Simulation: Lightweight simulator based on published ALFWorld information (papers, ReAct prompt examples).
Subgoal Decomposition: Rule-based… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/sft_alfworld_trajectory_dataset_v4.sft_alfworld_trajectory_dataset
ALFWorld Trajectory Dataset
Overview
This is a synthetic SFT (Supervised Fine-Tuning) dataset designed for agent training in ALFWorld-compatible environments. The dataset programmatically generates expert trajectories without requiring an actual ALFWorld environment or a large language model.
Key Approach
Template-based Simulation: Lightweight simulator based on published ALFWorld information (papers, ReAct prompt examples).
Subgoal Decomposition: Rule-based… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/sft_alfworld_trajectory_dataset.cot-trajectory-prompt-datasetv2agent-trajectory-dataso101_glue_dataset-trajectoryThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101",
"total_episodes": 51,
"total_frames": 22488,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/so101_glue_dataset-trajectory.sft_alfworld_trajectory_dataset_v3
ALFWorld Trajectory Dataset
Overview
This is a synthetic SFT (Supervised Fine-Tuning) dataset designed for agent training in ALFWorld-compatible environments. The dataset programmatically generates expert trajectories without requiring an actual ALFWorld environment or a large language model.
Key Approach
Template-based Simulation: Lightweight simulator based on published ALFWorld information (papers, ReAct prompt examples).
Subgoal Decomposition: Rule-based… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/sft_alfworld_trajectory_dataset_v3.sft_alfworld_trajectory_dataset_v2
ALFWorld Trajectory Dataset v1 (v5_cleaned + additive)
Overview
This dataset is a simple concatenation of:
Base: moroqq/sft_alfworld_trajectory_dataset_v5_cleaned
https://huggingface.co/datasets/moroqq/sft_alfworld_trajectory_dataset_v5_cleaned
Additive (local JSONL): additive_data/sft_alfworld_trajectory_additive_data_96cases.jsonl
The additive JSONL is normalized to match the base dataset schema (messages, metadata).
Statistics
base rows: 2277… See the full description on the dataset page: https://huggingface.co/datasets/moroqq/sft_alfworld_trajectory_dataset_v2.EEG-fNIRS-based-Handwriting-Trajectory-Dataset
Overview
This dataset contains the raw multimodal signals and metadata.
The goal of the competition is to predict the imagined handwriting class for each test trial using synchronized:
EEG signals
fNIRS signals
The dataset is organized to support an end-to-end competition workflow:
train_meta.csv provides labeled training trials
test_meta.csv provides unlabeled test trials
raw/ contains the corresponding raw EEG and fNIRS recordings
Files… See the full description on the dataset page: https://huggingface.co/datasets/lasfk/EEG-fNIRS-based-Handwriting-Trajectory-Dataset.sft_alfworld_trajectory_dataset_v5_cleaned
ALFWorld Trajectory Dataset v5 (Cleaned)
This dataset is a cleaned derivative of:
u-10bei/sft_alfworld_trajectory_dataset_v5
https://huggingface.co/datasets/u-10bei/sft_alfworld_trajectory_dataset_v5
What was removed?
We removed samples that contain the string Nothing happens. (case-insensitive) in any user message.
Rationale: those trajectories typically correspond to invalid / non-admissible actions (e.g. hallucinated object ids)
and can increase the probability of… See the full description on the dataset page: https://huggingface.co/datasets/moroqq/sft_alfworld_trajectory_dataset_v5_cleaned.sft_alfworld_trajectory_dataset_v3to5_admissible
Dataset Card for ALFWorld SFT Trajectories with Admissible Actions (v3–v5)
Overview
This dataset contains a subset of ALFWorld-style trajectories used for
supervised fine-tuning (SFT) of an agent that interacts with a textual
household environment.
It is constructed by merging v3, v4, and v5 of the original
u-10bei/sft_alfworld_trajectory_dataset_* series and filtering only
trajectories that include admissible actions in the observation text.
Each example corresponds to… See the full description on the dataset page: https://huggingface.co/datasets/kuririrn/sft_alfworld_trajectory_dataset_v3to5_admissible.sft_alfworld_trajectory_dataset_v3to5_admissible_success
Dataset Card for ALFWorld SFT Trajectories with Admissible Actions (v3–v5)
Overview
This dataset contains a subset of ALFWorld-style trajectories used for supervised fine-tuning (SFT) of an agent interacting with a textual household environment.
It is constructed by merging v3, v4, and v5 of the original
u-10bei/sft_alfworld_trajectory_dataset_* series and filtering only trajectories that include admissible actions in the observation text.
Each example corresponds to one… See the full description on the dataset page: https://huggingface.co/datasets/kuririrn/sft_alfworld_trajectory_dataset_v3to5_admissible_success.
