LLM agent
gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-heretic-GGUFllmfan46-gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-heretic-ROCMFPXnesso-0.4B-agenticgemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-hereticllm-agents_tora-code-7b-v1.0-GGUFtora-code-7b-v1.0Fabliq-8B-Agent-Mega-10ep-GGUFFabliq-8B-Agent-FromBase-Reasoning-GGUF
agent-llm-traces-v2
Exgentic Agent LLM Traces v2 — Agent Chat Only
OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it.
This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.LLM-Agent-Harness-Survey
English | 中文
Agent Harness for Large Language Model Agents: A Survey
⭐ This repo is actively maintained. If you find it useful, please star the repo to stay updated and help others find it.
The agent execution harness — not the model — is the primary determinant of agent reliability at scale.This survey formalizes the harness as a first-class architectural object H = (E, T, C, S, L, V), surveys 110+ papers, blogs and reports across 23 systems, and maps 9 open… See the full description on the dataset page: https://huggingface.co/datasets/GloriaaaM/LLM-Agent-Harness-Survey.agent-llm-traces
Multi-Benchmark LLM Agent Traces
A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization.
Collected by Exgentic - A platform for LLM observability and performance optimization.
Dataset Overview
This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces.automl_llm_agent_m2
AutoML-LLM Agent Module 2 Benchmark
This dataset contains the Module 2 benchmark for evaluating an assistant that converts a Module 1 recipe, a user request, and a processed tabular dataset into an auditable autogluon.cloud.TabularCloudPredictor configuration.
The repository is scoped to Module 2 only.
Tables
module2_cases: one row per Module 2 evaluation case.
module2_queries: user requests for each case.
module2_reference_configs: legacy reference JSON files… See the full description on the dataset page: https://huggingface.co/datasets/tecnologiactc/automl_llm_agent_m2.automl_llm_agent_m1
AutoML-LLM Agent Module 1 Benchmark
This dataset contains the Module 1 benchmark for evaluating an AutoML assistant that interprets user requests, selects tabular modeling settings, produces an auditable AutoGluon Tabular plan, and synthesizes the compact Module 1 recipe consumed by Module 2 through the mandatory final LLM writer used by all A-E variants.
The repository is scoped to Module 1 only.
Tables
cases: one row per Module 1 evaluation case.
queries: one… See the full description on the dataset page: https://huggingface.co/datasets/tecnologiactc/automl_llm_agent_m1.details_llm-agents__tora-70b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-70b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-70b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-70b-v1.0.
