datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Workspace-Bench
Workspace-Bench
Workspace-Bench is a benchmark for evaluating AI agents on realistic workspace tasks with large-scale file dependencies. It is designed to measure whether an agent can discover, interpret, and use the right files inside a noisy professional workspace, rather than solving tasks from isolated inputs or fully pre-packaged evidence.
The benchmark is introduced in the paper "Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File… See the full description on the dataset page: https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench.Workspace-Bench-Lite
Workspace-Bench-Lite
A Lightweight Subset of Workspace-Bench for Fast and Cost-Efficient Evaluation
Overview •
LeaderBoard •
Distribution •
Quick Start •
Changelog •
Citation
Overview
Workspace-Bench-Lite is the Lite split of Workspace-Bench 1.0, designed for fast iteration and lower-cost benchmarking while preserving the core evaluation setting of the full benchmark.
It contains 100 tasks selected from the full Workspace-Bench and is intended to… See the full description on the dataset page: https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench-Lite.kitchen-workspace-understanding-safe-manipulation
Kitchen Workspace Understanding & Safe Manipulation
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/kitchen-workspace-understanding-safe-manipulation.tinygsm_fobinary_workspace_depth1to9_traindepth5india-e1-workspace-mirrordedup-workspacetinygsm_fopython_workspace_depth1to9_traindepth5kaggle-workspace
BDMCPF
BD motorcycle ride POV footage, collected to experiment with road condition analysis from vehicle POV footage and road condition mapping.
Dataset Description
This dataset contains first-person (POV) video recordings of motorcycle rides in Bangladesh. The footage is intended for research on analyzing road conditions directly from vehicle POV video.
Use Cases
Road condition classification and analysis
Road surface quality mapping… See the full description on the dataset page: https://huggingface.co/datasets/arissassina/kaggle-workspace.zarn-workspace-rag-qa
Zarn Workspace RAG QA
Dataset Description
Small document bundles paired with grounded answers, evidence, and explicit refusals when context is missing.
Team Attribution
This dataset was created and reviewed by the Zarnite team through internal benchmark design, generation, and quality-control workflows. It should be presented as a Zarnite-authored benchmark starter pack, not as a purely human-collected field corpus.
Ecosystem Need Tier
High Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/zarnite/zarn-workspace-rag-qa.algo-sft-eval-traces-long-arithmetic-distill-qwq-v4
algo-sft-eval-traces-long-arithmetic-distill-qwq-v4
Full eval traces for algo-sft-long-arithmetic-distill-qwq across test/harder/ood splits
Dataset Info
Rows: 2000
Columns: 11
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier from eval set
split
Value('string')
Evaluation split: test (in-distribution), harder (scaled up), ood (structural out-of-distribution)
domain
Value('string')
Task domain: formal_logic… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algo-sft-eval-traces-long-arithmetic-distill-qwq-v4.algo-sft-eval-traces-conlang-morphology-ordered-rules-d5d7-v4
algo-sft-eval-traces-conlang-morphology-ordered-rules-d5d7-v4
Full eval traces for algo-sft-conlang-morphology-ordered-rules-d5d7 across test/harder/ood splits
Dataset Info
Rows: 2000
Columns: 11
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier from eval set
split
Value('string')
Evaluation split: test (in-distribution), harder (scaled up), ood (structural out-of-distribution)
domain
Value('string')
Task… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algo-sft-eval-traces-conlang-morphology-ordered-rules-d5d7-v4.prefvla-workspace-0611autotrainer-v0
autotrainer-v0
AutoTrainer-v0: LLM-agent-controlled GRPO training on Countdown
Dataset Info
Rows: 1
Columns: 1
Columns
Column
Type
Description
state_json
Value('string')
Full autotrainer state as JSON string
Generation Parameters
{
"script_name": "run_round.py",
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"description": "AutoTrainer-v0: LLM-agent-controlled GRPO training on Countdown",
"experiment_id": "autotrainer-v0"… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/autotrainer-v0.dspy-security-bench-trainset-workspace
dspy-security-bench: workspace trainset (v0.1)
This is the synthetic, environment-grounded query-only trainset used to
optimize DSPy programs in v0.1 of
dspy-security-bench,
a benchmark that measures whether DSPy prompt optimization affects the
prompt-injection robustness of agentic LLM programs.
What's in here
192 query / ground-truth pairs grounded in the
AgentDojo workspace suite's
default environment (calendar, inbox, files).
{"prompt": "What is the… See the full description on the dataset page: https://huggingface.co/datasets/immu4989/dspy-security-bench-trainset-workspace.google-workspace-dataset
google-workspace-dataset
Dataset generated with DeepFabric.
algo-sft-eval-traces-cellular-automata-step-simulation-d5-v4
algo-sft-eval-traces-cellular-automata-step-simulation-d5-v4
Full eval traces for algo-sft-cellular-automata-step-simulation-d5 across test/harder/ood splits
Dataset Info
Rows: 2000
Columns: 11
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier from eval set
split
Value('string')
Evaluation split: test (in-distribution), harder (scaled up), ood (structural out-of-distribution)
domain
Value('string')
Task… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algo-sft-eval-traces-cellular-automata-step-simulation-d5-v4.workspace
wdv3-timm
small example thing showing how to use timm to run the WD Tagger V3 models.
How To Use
clone the repository and enter the directory:
git clone https://github.com/neggles/wdv3-timm.git
cd wd3-timm
Create a virtual environment and install the Python requirements.
If you're using Linux, you can use the provided script:
bash setup.sh
Or if you're on Windows (or just want to do it manually), you can do the following:
# Create virtual environment
python3.10 -m… See the full description on the dataset page: https://huggingface.co/datasets/codeShare/workspace.algo-sft-eval-baseline-Qwen2.5-1.5B-Instruct-v4
algo-sft-eval-baseline-Qwen2.5-1.5B-Instruct-v4
Baseline eval traces for Qwen2.5-1.5B-Instruct across all 4 domains × 3 splits
Dataset Info
Rows: 8000
Columns: 10
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier from eval set
domain
Value('string')
Task domain
split
Value('string')
Evaluation split: test, harder, ood
prompt
Value('string')
Full prompt sent to the model
model_response
Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algo-sft-eval-baseline-Qwen2.5-1.5B-Instruct-v4.autotrainer-v1
AutoTrainer v1
Agentic training harness where Claude decides every training step's method, data, and hyperparameters.
Configs
steps: Per-step decisions, metrics, agent traces, and costs
eval_traces: Per-question evaluation results with difficulty breakdown (n_args=2-10)
autotrainer-v1-run-v2-11stepsalgo-sft-eval-traces-formal-logic-bottom-up-v4
algo-sft-eval-traces-formal-logic-bottom-up-v4
Full eval traces for algo-sft-formal-logic-bottom-up across test/harder/ood splits
Dataset Info
Rows: 2000
Columns: 11
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier from eval set
split
Value('string')
Evaluation split: test (in-distribution), harder (scaled up), ood (structural out-of-distribution)
domain
Value('string')
Task domain: formal_logic… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algo-sft-eval-traces-formal-logic-bottom-up-v4.PROJECT-MANIFEST
PROJECT-MANIFEST
Central registry of all datasets in the raca-workspace-v1 organization.
Total Datasets Tracked: 17
Last Updated: 2026-04-19T07:30:07.171666+00:00
Usage
from datasets import load_dataset
manifest = load_dataset("raca-workspace-v1/PROJECT-MANIFEST", split="train")
print(f"Tracking {len(manifest)} datasets")
Automatically managed by RACA hf_utility.
sdc-all-responses-v1
sdc-all-responses-v1
Complete semantic-distance-coding experiment: 15 programming languages × 4 difficulty tiers × 20 EsoLang-Bench problems × 3 runs. Full prompts, model responses, extracted code, compilation results, and per-test-case outcomes. Wave 1 (8 mainstream: Python, C++, Java, Perl, Rust, Go, Haskell, OCaml) + Wave 2 (7 added: Fortran, Ada, Prolog, COBOL, F#, Erlang, Tcl). Zero-shot condition with GPT-5.2.
Dataset Info
Rows: 3600
Columns: 16
Columns… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/sdc-all-responses-v1.algo-sft-eval-traces-conlang-morphology-distill-qwq-v4
algo-sft-eval-traces-conlang-morphology-distill-qwq-v4
Full eval traces for algo-sft-conlang-morphology-distill-qwq across test/harder/ood splits
Dataset Info
Rows: 2000
Columns: 11
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier from eval set
split
Value('string')
Evaluation split: test (in-distribution), harder (scaled up), ood (structural out-of-distribution)
domain
Value('string')
Task domain:… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algo-sft-eval-traces-conlang-morphology-distill-qwq-v4.algo-sft-eval-traces-cellular-automata-distill-qwq-v4
algo-sft-eval-traces-cellular-automata-distill-qwq-v4
Full eval traces for algo-sft-cellular-automata-distill-qwq across test/harder/ood splits
Dataset Info
Rows: 2000
Columns: 11
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier from eval set
split
Value('string')
Evaluation split: test (in-distribution), harder (scaled up), ood (structural out-of-distribution)
domain
Value('string')
Task domain:… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algo-sft-eval-traces-cellular-automata-distill-qwq-v4.algorithmic-sft-full-eval-v4
algorithmic-sft-full-eval-v4
Aggregate eval results: 10 models x 4 domains x 3 splits with bootstrap 95% CIs
Dataset Info
Rows: 42
Columns: 8
Columns
Column
Type
Description
model
Value('string')
HuggingFace model ID (LoRA adapter name)
domain
Value('string')
Task domain: formal_logic, conlang_morphology, cellular_automata, long_arithmetic
type
Value('string')
Training type: algo (algorithmic SFT) or distill (QwQ distillation)
split… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algorithmic-sft-full-eval-v4.ttt-discover-circle_packing_26-qwen3-8b-v1grpo-tool-sat-dataset-v1
grpo-tool-sat-dataset-v1
Synthetic lookup-table dataset for the GRPO Tool Saturation experiment. 10k keys k in [0, 9999] with r = k mod 3 feature. Tools map/table have opaque, partially overlapping correct-domains: map correct for r in {0, 2}, table correct for r in {1, 2}. f(k) = SHA256(str(k))[:6]; wrong-hash returns are g_map(k), g_table(k). SFT demos use overlap_skew_map=0.6 on r=2; token-level shuffle across r-classes with fixed seed.
Dataset Info
Rows: 20000… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/grpo-tool-sat-dataset-v1.grpo-tool-sat-sft-corpus-v1
grpo-tool-sat-sft-corpus-v1
SFT-only view of grpo-tool-sat-dataset-v1. v1.1 — prompt now ends with newline to match RL eval_prefix concatenation.
Dataset Info
Rows: 8000
Columns: 5
Columns
Column
Type
Description
k
Value('int64')
Key integer
r
Value('int64')
k mod 3
tool
Value('string')
map or table
prompt
Value('string')
User prompt: "Key:
" (trailing newline)
response
Value('string')
Target completion: prose + + +… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/grpo-tool-sat-sft-corpus-v1.humanoid-workspace-organization-tr-v1Workspace organization dataset for humanoid robots.
Description
Indoor desk organization and task finalization interactions.
Task Description
Teaches humanoid robots to organize workspaces, manage office items and ensure safe shutdown procedures after task completion.
