CoolFace
Datasetpublic

antoinezambelli/forge-evals

Forge Agentic Workflow Evaluation Corpus Forge evaluates multi-step agent workflows across models, quantizations, inference backends, function-calling modes, scenarios, guardrail ablations, and reasoning-replay policies. This dataset publishes the released run-level outcome records behind Forge's reports and dashboard. One row represents one attempted evaluation run. This is an outcome corpus, not a collection of complete agent traces. It does not contain full prompts… See the full description on the dataset page: https://huggingface.co/datasets/antoinezambelli/forge-evals.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
1likes70downloads

antoinezambelli/forge-evals · main · files are served by the source, never re-hosted here