Rislantrs/TelAgentBench-ID
TelAgentBench-ID: A Comprehensive Benchmark for Evaluating Autonomous LLM Agents in Telecommunications Business Support Systems ๐ Dataset Summary TelAgentBench-ID is the first comprehensive, multi-faceted benchmark specifically constructed to evaluate the Action Execution Fidelity and Epistemic Calibration of Large Language Models (LLMs) and Small Language Models (SLMs) within the Telecommunications Business Support Systems (BSS) domain in Indonesian.โฆ See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelAgentBench-ID.
TelAgentBench-ID: A Comprehensive Benchmark for Evaluating Autonomous LLM Agents in Telecommunications Business Support Systems
   
๐ Dataset Summary
TelAgentBench-ID is the first comprehensive, multi-faceted benchmark specifically constructed to evaluate the Action Execution Fidelity and Epistemic Calibration of Large Language Models (LLMs) and Small Language Models (SLMs) within the Telecommunications Business Support Systems (BSS) domain in Indonesian.
Adapted and localized from the international benchmark TelAgentBench (Lee et al., EMNLP 2025 Industry Track), TelAgentBench-ID standardizes 23 realistic enterprise BSS API endpoints based on GSMA Open Gateway and TM Forum Open Digital Architecture (ODA) specifications, mapped to the Berkeley Function Calling Leaderboard (BFCL) evaluation format.
The benchmark spans 757 test scenarios in tool-calling action evaluation (complemented by instruction-following and multi-city travel planning modules, totaling 2,137 evaluation instances across 40 JSON/JSONL assets). It is explicitly engineered to diagnose both Execution Fidelity (strict parameter extraction and syntactic JSON compliance) and Epistemic Calibration (mitigation of Compulsive Action Bias / Over-Triggering on ambiguous or incomplete transactions).
๐๏ธ Background & Motivation
In mission-critical enterprise environments like telecommunications BSS (governing billing settlement, real-time quota deduction, and supplementary add-on management):
- Customer Data Sovereignty: Under statutes such as the Indonesian Personal Data Protection Law (UU PDP No. 27/2022) and the GDPR, cross-border transmission of sensitive telecommunications payloads (MSISDN, billing records, identity tokens) to multi-tenant public cloud APIs is legally constrained. Evaluating parameter-efficient Small Language Models (SLMs, $\le$ 7B) capable of on-premise deployment is paramount.
- Deterministic Execution vs. Free-Form Generation: Telecommunications transactions necessitate strictly valid, deterministic API payloads. Textual generation metrics (such as BLEU or ROUGE) are inadequate for measuring functional correctness.
- Epistemic Calibration: Production agents must strike a calibrated equilibrium: executing tools with 0.00% premature refusal on valid queries, while actively refraining (calibrated abstention) when customer inputs lack necessary credentials, preventing database corruption.
๐งฑ Benchmark Architecture & API Testbed
The evaluation suite models 23 production-grade BSS API endpoints categorized across six functional service clusters:
๐ Directory Structure & Modularity
The benchmark repository contains 40 verified, valid JSON/JSONL evaluation assets:
TelAgentBench-ID/
โโโ README.md # Comprehensive Dataset Card
โโโ validation.json # Multi-City Travel Planning & Roaming (200 instances)
โโโ TelAgent_Plan/
โ โโโ validation_dataset_1111.json # Detailed Plan Evaluation Corpus (200 instances)
โโโ TelAgent_IF/
โ โโโ telif_general_ko.json # Multi-Turn Instruction Following (200 instances, 3-turn)
โโโ TelAgent_Action/
โโโ 250918_telcoFC_tools.json # Master Tool Schemas (23 BSS APIs)
โโโ telco_simple.json # Single-Turn 1-Candidate Tool Evaluation (55 items)
โโโ telco_multiple.json # Single-Turn 7-Candidate Tool Evaluation (55 items)
โโโ telco_parallel.json # Parallel Multi-Tool Invocations (19 items)
โโโ telco_parallel_multiple.json # Parallel Multi-Tool with Distractor Schemas (18 items)
โโโ telco_item_recommendation_parallel_multiple.json # Multi-Constraint Plan Recommendations (26 items)
โโโ telco_multi_step_stateless_partial.json # Multi-Step Tool Execution with Partial Schemas (23 items)
โโโ telco_multi_step_stateless_whole.json # Multi-Step Tool Execution with Full Schemas (23 items)
โโโ telco_multi_turn_stateless_base.json # Conversational Multi-Turn State Tracking (46 items, 3 turns)
โโโ telco_multi_turn_stateless_miss_param.json # Conversational Multi-Turn Missing Parameter Slot-Filling (38 items)
โโโ telco_live_simple.json # Live Persona-Prompted Single-Turn Simple (55 items)
โโโ telco_live_multiple.json # Live Persona-Prompted Single-Turn Multiple (55 items)
โโโ telco_live_parallel.json # Live Persona-Prompted Parallel Execution (19 items)
โโโ telco_live_parallel_multiple.json # Live Persona-Prompted Parallel Multi-Tool (18 items)
โโโ telco_live_item_recommendation_parallel_multiple.json # Live Recommendation Evaluation (26 items)
โโโ telco_live_multi_step_stateless_partial.json # Live Multi-Step Execution Partial (23 items)
โโโ telco_live_multi_step_stateless_whole.json # Live Multi-Step Execution Whole (23 items)
โโโ telco_live_relevance.json # Live Relevance Discrimination (117 items)
โโโ telco_live_irrelevance.json # Calibrated Abstention / Unanswerable Queries (118 items)
โโโ possible_answer/ # 18 Matching Ground-Truth AST Formatted Files
โโโ telco_simple.json
โโโ telco_multiple.json
โโโ telco_multi_turn_stateless_base.json
โโโ ... (15 other ground-truth files)๐งช Evaluation Metrics & Protocols
All models evaluated against TelAgentBench-ID are assessed across standardized AST execution metrics:
- Tool Selection Accuracy: $$\text{Acc}{\text{tool}} = \frac{\sum \mathbb{I}(f{\text{pred}} = f_{\text{gt}})}{N}$$ Measures intent recognition fidelity (identifying the correct target endpoint among candidates).
- Strict AST Accuracy: $$\text{Acc}{\text{AST}} = \frac{\sum \mathbb{I}(f{\text{pred}} = f{\text{gt}} \land \Theta{\text{pred}} \equiv \Theta_{\text{gt}})}{N}$$ Evaluates full functional determinism: verifies that function name, argument keys, data types, and extracted parameter values (MSISDNs, formatted date ranges, product IDs) match the ground-truth Abstract Syntax Tree exactly.
- Format Validity: $$\text{Acc}{\text{format}} = \frac{\sum \mathbb{I}(\text{is\valid\_json}(y))}{N}$$ Measures the structural compliance of model outputs (100.00% indicates zero JSON parser crashes).
- Epistemic Calibration Metrics:
- Premature Abstention: False refusal rate on fully-specified, answerable queries (ideal target: $0.00\%$).
- Calibrated Abstention: Correct refusal rate on underspecified, ambiguous, or out-of-domain queries (suppressing Compulsive Action Bias).
๐ฌ Benchmark Baseline Results
Benchmark results from the accompanying study (Tristansyah & Ichsan, 2026) under deterministic decoding ($\text{Temperature} = 0.0$):
Key Insights:
- Domain Weight Specialization Outperforms Scale:
TelkomNusa-LLM-7Bachieves 83.64% Single-Turn AST, surpassing the massive GPT-OSS-120B (80.00%). - Edge Viability:
TelkomNusa-SLM-3Bachieves 44.55% Overall AST, outperformingLlama-3.2-3B(41.09%) with over $53\%$ lower latency on a commodity NVIDIA Tesla T4 GPU.
๐ Data Hygiene & Contamination Safeguards
To prevent artificial metric inflation and data leakage:
- Strictly Disjoint Testbed: The 757 evaluation scenarios are strictly quarantined (held-out) from the 3,510 training pairs of the TelkomNusa SFT corpus.
- Decontamination Protocol: N-gram overlap and AST similarity audits were executed following the contamination protocols established by Deng et al. (NAACL 2024) and Lee et al. (ACL 2022).
- Synthetic PII Masking: All MSISDNs utilize synthetic Indonesian telecommunications number allocations (
0812xxxx,0811xxxx). No real customer Personally Identifiable Information (PII) is included, ensuring compliance with Indonesian Law No. 27/2022 (UU PDP).
๐ How to Cite
If you utilize TelAgentBench-ID in your research, please cite both the accompanying study and the original foundational benchmark:
@article{tristansyah2026benchmarking,
title={Benchmarking Small Language Models Against Frontier Giants for Telecommunications BSS Tool-Calling: Execution Fidelity and Epistemic Calibration},
author={Tristansyah, M Rislan and Ichsan, Ichwan Nul},
journal={JUTI: Jurnal Ilmiah Teknologi Informasi},
volume={x},
number={x},
year={2026},
publisher={Universitas Pendidikan Indonesia \& Institut Teknologi Sepuluh Nopember}
}
@inproceedings{lee2025telagentbench,
title={TelAgentBench: A Multi-Faceted Benchmark for Evaluating LLM-Based Agents in Telecommunications},
author={Lee, S. and Potdar, S. and Rojas-Barahona, L. and Montella, S. and others},
booktitle={Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP): Industry Track},
pages={1173--1211},
year={2025},
publisher={Association for Computational Linguistics},
doi={10.18653/v1/2025.emnlp-industry.83}
}๐ License & Attribution
This dataset is distributed under the Creative Commons Attribution 4.0 International License ([CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)). You are free to share, copy, modify, and build upon this benchmark for academic and commercial purposes, provided appropriate attribution is given to the original authors.
