agent-benchmarks
data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.windows-agent-benchmarks
Windows Agent Benchmark
Evaluates models on Windows-specific PowerShell tasks across 4 domains:
Windows Auto: Service management, scheduled tasks, registry, DISM, event logs, WinRM
Active Directory / Entra ID: User/group/OU management, Graph PowerShell, AAD Connect, Conditional Access
Intune / MECM: Device management, compliance, Autopilot, BitLocker, ConfigMgr deployments
M365 / Exchange: Mailboxes, Teams, SharePoint, Purview DLP, transport rules, hybrid… See the full description on the dataset page: https://huggingface.co/datasets/Ianinh0/windows-agent-benchmarks.ai-agent-benchmarks
AI Agent Benchmarks Dataset
Benchmark data for AI agent performance across various tasks.
About
This dataset contains benchmarks for AI agents including:
Task completion rates
Response times
Accuracy metrics
Multi-step reasoning performance
Source
Collected by Creative Content Crafts for the Co.Actor platform.
Related Resources
Co.Actor: https://co.actor - AI collaborative automation
Company: Creative Content Crafts
Wikidata: Q137625544… See the full description on the dataset page: https://huggingface.co/datasets/sergeinboca/ai-agent-benchmarks.p2pclaw-agent-benchmarks
