penfever/large-model-agentic-eval-parity-tb21-aa
Large Model Agentic Eval Parity with AAII on TB2.1 This repository contains the launch configurations, Harbor results, trajectories, and analysis artifacts for marin-community/marin#8261. The campaign evaluates Qwen3.5-122B-A10B-FP8, DeepSeek-V4-Flash-0731, Nemotron 3 Ultra 550B A55B NVFP4, and GLM-5.2 AWQ INT4 against Artificial Analysis Terminal-Bench 2.1 results. Redaction Resolved Harbor records and captured terminal output contained signed endpoint URLs and… See the full description on the dataset page: https://huggingface.co/datasets/penfever/large-model-agentic-eval-parity-tb21-aa.
This repository belongs to penfever on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
