CoolFace
Datasetpublic

antoinezambelli/forge-evals

Forge Agentic Workflow Evaluation Corpus Forge evaluates multi-step agent workflows across models, quantizations, inference backends, function-calling modes, scenarios, guardrail ablations, and reasoning-replay policies. This dataset publishes the released run-level outcome records behind Forge's reports and dashboard. One row represents one attempted evaluation run. This is an outcome corpus, not a collection of complete agent traces. It does not contain full prompts… See the full description on the dataset page: https://huggingface.co/datasets/antoinezambelli/forge-evals.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
1likes70downloads
5 commits on main
4e805f01mo ago

Update Forge eval dataset (#4)

antoinezambelli
35f2cd91mo ago

Update Forge eval dataset with Inkling Small sweep (#3)

antoinezambelli
eec38431mo ago

Publish expanded Forge evaluation corpus (#2)

antoinezambelli
6a02fc32mo ago

Publish Forge evaluation corpus v1

antoinezambelli
c1e2f2d2mo ago

initial commit

antoinezambelli