penfever/large-model-agentic-eval-parity-tb21-aa
Large Model Agentic Eval Parity with AAII on TB2.1 This repository contains the launch configurations, Harbor results, trajectories, and analysis artifacts for marin-community/marin#8261. The campaign evaluates Qwen3.5-122B-A10B-FP8, DeepSeek-V4-Flash-0731, Nemotron 3 Ultra 550B A55B NVFP4, and GLM-5.2 AWQ INT4 against Artificial Analysis Terminal-Bench 2.1 results. Redaction Resolved Harbor records and captured terminal output contained signed endpoint URLs and… See the full description on the dataset page: https://huggingface.co/datasets/penfever/large-model-agentic-eval-parity-tb21-aa.
Large Model Agentic Eval Parity with AAII on TB2.1
This repository contains the launch configurations, Harbor results, trajectories, and analysis artifacts for marin-community/marin#8261.
The campaign evaluates Qwen3.5-122B-A10B-FP8, DeepSeek-V4-Flash-0731, Nemotron 3 Ultra 550B A55B NVFP4, and GLM-5.2 AWQ INT4 against Artificial Analysis Terminal-Bench 2.1 results.
Redaction
Resolved Harbor records and captured terminal output contained signed endpoint URLs and transient credentials. Public copies preserve the original paths while replacing credential material with <redacted>. `REDACTED_FILES.txt` lists every affected file.
