automation
pipeline_automation
Circuit Tracing Automation: LLMs can annotate attribution graphs
Custom automation pipeline on top of the circuit-tracer library. Automatically generates feature descriptions, supernodes, and validation scores for an attribution graph.
Repository Structure
prompts/ # Prompt datasets (shared across models)
prompts_capital.csv # Input prompts for attribution
ground_truth_capital.csv # Correct answers + intermediate… See the full description on the dataset page: https://huggingface.co/datasets/circuit-tracer-automation/pipeline_automation.automationbench600-open-weight-runs
AutomationBench 600 — raw run artifacts (open-weight models)
Complete raw exports from running the public 600-task
AutomationBench (Zapier, v1.0.5, --toolset api)
on open-weight models served locally with SGLang 0.5.12 on 8x H200.
Each .json.gz is the untouched --export-json output: run metadata, summary, and one record per
task including every message of the trajectory, the final environment state, and per-assertion
results.
File
Model
Pass rate
Partial credit… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/automationbench600-open-weight-runs.bench-automationbench
AutomationBench public task payloads
Materialized public task inputs for Zapier AutomationBench,
pinned to upstream revision 4a8e1061254004d9dac807054eed33fad7d1ff14.
This repository contains six Parquet files (100 tasks per public domain) generated from the
upstream get_<domain>_dataset() functions. It exists so the solar-system evaluation integration
can fetch immutable task payloads without committing multi-megabyte generated Python modules.
Source and license: Zapier… See the full description on the dataset page: https://huggingface.co/datasets/hyeonseop-upstage/bench-automationbench.lerobot-record-sofficebench-automation
officebench-automation
OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
Dataset Description
This dataset is a standardized version of the original benchmark, prepared for easy evaluation of LLMs on planning tasks.
Splits
test: 300 samples
Features
{
"id": "Value(dtype='string', id=None)",
"task_id": "Value(dtype='string', id=None)",
"subtask_id": "Value(dtype='string', id=None)",
"num_apps":… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/officebench-automation.Octpickandplace
Octpickandplace
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
