datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.windows-agent-benchmarks
Windows Agent Benchmark
Evaluates models on Windows-specific PowerShell tasks across 4 domains:
Windows Auto: Service management, scheduled tasks, registry, DISM, event logs, WinRM
Active Directory / Entra ID: User/group/OU management, Graph PowerShell, AAD Connect, Conditional Access
Intune / MECM: Device management, compliance, Autopilot, BitLocker, ConfigMgr deployments
M365 / Exchange: Mailboxes, Teams, SharePoint, Purview DLP, transport rules, hybrid… See the full description on the dataset page: https://huggingface.co/datasets/Ianinh0/windows-agent-benchmarks.
