workspace
Workspace-Bench-Workspaces
Workspace-Bench-Workspaces
Workspace-Bench-Workspaces contains the large workspace filesystem archives used by Workspace-Bench and Workspace-Bench-Lite. These files provide the realistic directory trees that agents must explore when solving workspace tasks.
The task metadata is hosted separately:
Full task metadata: Workspace-Bench
Lite task metadata: Workspace-Bench-Lite
Contents
This repository contains two workspace archives:
filesys_en.zip: English… See the full description on the dataset page: https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench-Workspaces.Workspace-Bench
Workspace-Bench
Workspace-Bench is a benchmark for evaluating AI agents on realistic workspace tasks with large-scale file dependencies. It is designed to measure whether an agent can discover, interpret, and use the right files inside a noisy professional workspace, rather than solving tasks from isolated inputs or fully pre-packaged evidence.
The benchmark is introduced in the paper "Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File… See the full description on the dataset page: https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench.Workspace-Bench-Lite
Workspace-Bench-Lite
A Lightweight Subset of Workspace-Bench for Fast and Cost-Efficient Evaluation
Overview •
LeaderBoard •
Distribution •
Quick Start •
Changelog •
Citation
Overview
Workspace-Bench-Lite is the Lite split of Workspace-Bench 1.0, designed for fast iteration and lower-cost benchmarking while preserving the core evaluation setting of the full benchmark.
It contains 100 tasks selected from the full Workspace-Bench and is intended to… See the full description on the dataset page: https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench-Lite.kitchen-workspace-understanding-safe-manipulation
Kitchen Workspace Understanding & Safe Manipulation
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/kitchen-workspace-understanding-safe-manipulation.local-workspace
local-workspace — consolidated oracle-lens artifacts
Everything the oracle-lens (OLA) program published, in one repo. Consolidated 2026-08-06 from
three repos (oracle-lens-data, oracle-lens-ar-checkpoints, oracle-lens-ao-checkpoints)
by scripts/ola/migrate_hf_workspace.py in the global-workspace repo, which also holds the
exact old->new path map and a --verify audit mode.
Layout
data/ # training/eval data (was oracle-lens-data)… See the full description on the dataset page: https://huggingface.co/datasets/agu18dec/local-workspace.vsi-workspace
VSI-Bench spatial-code pipeline
Video → 3D spatial reasoning: Depth Anything 3 (metric depth + camera poses) and SAM 3
(text-prompted instance segmentation) are fused into a per-scene spatial code — a JSON of
object positions, sizes, distances, room geometry, appearance order, and camera trajectory. A
unified vision-language model (Qwen3.5-2B / 4B) then answers VSI-Bench questions from one of
three inputs: the spatial code alone (code), 32 uniform video frames alone (frames), or… See the full description on the dataset page: https://huggingface.co/datasets/AntonioJun/vsi-workspace.
