CoolFace
Datasetpublic

Rislantrs/TelAgentBench-ID

TelAgentBench-ID: A Comprehensive Benchmark for Evaluating Autonomous LLM Agents in Telecommunications Business Support Systems ๐Ÿ“Œ Dataset Summary TelAgentBench-ID is the first comprehensive, multi-faceted benchmark specifically constructed to evaluate the Action Execution Fidelity and Epistemic Calibration of Large Language Models (LLMs) and Small Language Models (SLMs) within the Telecommunications Business Support Systems (BSS) domain in Indonesian.โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelAgentBench-ID.

sourceHugging Facecc-by-4.0updated 6d agoView on Hugging Face
0likes579downloads
Dataset Card

TelAgentBench-ID: A Comprehensive Benchmark for Evaluating Autonomous LLM Agents in Telecommunications Business Support Systems

![License: CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) ![Language: Indonesian-blue.svg)](https://huggingface.co/languages) ![Domain: Telecommunications BSS](https://www.gsma.com/solutions-and-impact/technologies/networks/gsma-open-gateway/) ![Framework: BFCL Protocol](https://gorilla.cs.berkeley.edu/leaderboard.html)


๐Ÿ“Œ Dataset Summary

TelAgentBench-ID is the first comprehensive, multi-faceted benchmark specifically constructed to evaluate the Action Execution Fidelity and Epistemic Calibration of Large Language Models (LLMs) and Small Language Models (SLMs) within the Telecommunications Business Support Systems (BSS) domain in Indonesian.

Adapted and localized from the international benchmark TelAgentBench (Lee et al., EMNLP 2025 Industry Track), TelAgentBench-ID standardizes 23 realistic enterprise BSS API endpoints based on GSMA Open Gateway and TM Forum Open Digital Architecture (ODA) specifications, mapped to the Berkeley Function Calling Leaderboard (BFCL) evaluation format.

The benchmark spans 757 test scenarios in tool-calling action evaluation (complemented by instruction-following and multi-city travel planning modules, totaling 2,137 evaluation instances across 40 JSON/JSONL assets). It is explicitly engineered to diagnose both Execution Fidelity (strict parameter extraction and syntactic JSON compliance) and Epistemic Calibration (mitigation of Compulsive Action Bias / Over-Triggering on ambiguous or incomplete transactions).


๐Ÿ›๏ธ Background & Motivation

In mission-critical enterprise environments like telecommunications BSS (governing billing settlement, real-time quota deduction, and supplementary add-on management):

  1. 1.Customer Data Sovereignty: Under statutes such as the Indonesian Personal Data Protection Law (UU PDP No. 27/2022) and the GDPR, cross-border transmission of sensitive telecommunications payloads (MSISDN, billing records, identity tokens) to multi-tenant public cloud APIs is legally constrained. Evaluating parameter-efficient Small Language Models (SLMs, $\le$ 7B) capable of on-premise deployment is paramount.
  2. 2.Deterministic Execution vs. Free-Form Generation: Telecommunications transactions necessitate strictly valid, deterministic API payloads. Textual generation metrics (such as BLEU or ROUGE) are inadequate for measuring functional correctness.
  3. 3.Epistemic Calibration: Production agents must strike a calibrated equilibrium: executing tools with 0.00% premature refusal on valid queries, while actively refraining (calibrated abstention) when customer inputs lack necessary credentials, preventing database corruption.

๐Ÿงฑ Benchmark Architecture & API Testbed

The evaluation suite models 23 production-grade BSS API endpoints categorized across six functional service clusters:

NoService ClusterEndpoints (APIs)Function ExamplesKey Parameters
1Bill Management3GET__rate_unpaid-bills, GET__rate_current-bills, GET__rate_fixed-billssvcMgmtNum (MSISDN), startDt, endDt
2Payment & Micro-Billing3GET__rate_content-purchases, GET__rate_mobile-payments, GET__rate_auto-paymentssvcMgmtNum, startDt, endDt
3Services & Add-ons5GET__rate_add-on-subscriptions, DELETE__rate_add-on-subscriptions, POST__call_spam-blocking, GET__call_coloringsvcMgmtNum, prodId
4Roaming Governance4GET__plan_roaming, POST__plan_roaming_subscriptions, DELETE__plan_roaming_subscriptions, GET__plan_roaming_subscriptionssvcMgmtNum, countryCode, countryName
5Refill & Data Coupons4GET__rate_refill-coupons, POST__rate_refill-coupons_register, POST__rate_refill-coupons_giftsvcMgmtNum, couponNum, receiverSvcMgmtNum
6Family BSS Plans4GET__rate_family_real-time-bill, GET__rate_family-representatives, GET__rate_family-network-groupssvcMgmtNum, familySvcMgmtNum

๐Ÿ“‚ Directory Structure & Modularity

The benchmark repository contains 40 verified, valid JSON/JSONL evaluation assets:

text
TelAgentBench-ID/
โ”œโ”€โ”€ README.md                                             # Comprehensive Dataset Card
โ”œโ”€โ”€ validation.json                                       # Multi-City Travel Planning & Roaming (200 instances)
โ”œโ”€โ”€ TelAgent_Plan/
โ”‚   โ””โ”€โ”€ validation_dataset_1111.json                      # Detailed Plan Evaluation Corpus (200 instances)
โ”œโ”€โ”€ TelAgent_IF/
โ”‚   โ””โ”€โ”€ telif_general_ko.json                             # Multi-Turn Instruction Following (200 instances, 3-turn)
โ””โ”€โ”€ TelAgent_Action/
    โ”œโ”€โ”€ 250918_telcoFC_tools.json                         # Master Tool Schemas (23 BSS APIs)
    โ”œโ”€โ”€ telco_simple.json                                 # Single-Turn 1-Candidate Tool Evaluation (55 items)
    โ”œโ”€โ”€ telco_multiple.json                               # Single-Turn 7-Candidate Tool Evaluation (55 items)
    โ”œโ”€โ”€ telco_parallel.json                               # Parallel Multi-Tool Invocations (19 items)
    โ”œโ”€โ”€ telco_parallel_multiple.json                      # Parallel Multi-Tool with Distractor Schemas (18 items)
    โ”œโ”€โ”€ telco_item_recommendation_parallel_multiple.json # Multi-Constraint Plan Recommendations (26 items)
    โ”œโ”€โ”€ telco_multi_step_stateless_partial.json           # Multi-Step Tool Execution with Partial Schemas (23 items)
    โ”œโ”€โ”€ telco_multi_step_stateless_whole.json             # Multi-Step Tool Execution with Full Schemas (23 items)
    โ”œโ”€โ”€ telco_multi_turn_stateless_base.json              # Conversational Multi-Turn State Tracking (46 items, 3 turns)
    โ”œโ”€โ”€ telco_multi_turn_stateless_miss_param.json        # Conversational Multi-Turn Missing Parameter Slot-Filling (38 items)
    โ”œโ”€โ”€ telco_live_simple.json                            # Live Persona-Prompted Single-Turn Simple (55 items)
    โ”œโ”€โ”€ telco_live_multiple.json                          # Live Persona-Prompted Single-Turn Multiple (55 items)
    โ”œโ”€โ”€ telco_live_parallel.json                          # Live Persona-Prompted Parallel Execution (19 items)
    โ”œโ”€โ”€ telco_live_parallel_multiple.json                 # Live Persona-Prompted Parallel Multi-Tool (18 items)
    โ”œโ”€โ”€ telco_live_item_recommendation_parallel_multiple.json # Live Recommendation Evaluation (26 items)
    โ”œโ”€โ”€ telco_live_multi_step_stateless_partial.json      # Live Multi-Step Execution Partial (23 items)
    โ”œโ”€โ”€ telco_live_multi_step_stateless_whole.json        # Live Multi-Step Execution Whole (23 items)
    โ”œโ”€โ”€ telco_live_relevance.json                         # Live Relevance Discrimination (117 items)
    โ”œโ”€โ”€ telco_live_irrelevance.json                       # Calibrated Abstention / Unanswerable Queries (118 items)
    โ””โ”€โ”€ possible_answer/                                  # 18 Matching Ground-Truth AST Formatted Files
        โ”œโ”€โ”€ telco_simple.json
        โ”œโ”€โ”€ telco_multiple.json
        โ”œโ”€โ”€ telco_multi_turn_stateless_base.json
        โ””โ”€โ”€ ... (15 other ground-truth files)

๐Ÿงช Evaluation Metrics & Protocols

All models evaluated against TelAgentBench-ID are assessed across standardized AST execution metrics:

  1. 1.Tool Selection Accuracy: $$\text{Acc}{\text{tool}} = \frac{\sum \mathbb{I}(f{\text{pred}} = f_{\text{gt}})}{N}$$ Measures intent recognition fidelity (identifying the correct target endpoint among candidates).
  2. 2.Strict AST Accuracy: $$\text{Acc}{\text{AST}} = \frac{\sum \mathbb{I}(f{\text{pred}} = f{\text{gt}} \land \Theta{\text{pred}} \equiv \Theta_{\text{gt}})}{N}$$ Evaluates full functional determinism: verifies that function name, argument keys, data types, and extracted parameter values (MSISDNs, formatted date ranges, product IDs) match the ground-truth Abstract Syntax Tree exactly.
  3. 3.Format Validity: $$\text{Acc}{\text{format}} = \frac{\sum \mathbb{I}(\text{is\valid\_json}(y))}{N}$$ Measures the structural compliance of model outputs (100.00% indicates zero JSON parser crashes).
  4. 4.Epistemic Calibration Metrics:
  5. 5.Premature Abstention: False refusal rate on fully-specified, answerable queries (ideal target: $0.00\%$).
  6. 6.Calibrated Abstention: Correct refusal rate on underspecified, ambiguous, or out-of-domain queries (suppressing Compulsive Action Bias).

๐Ÿ”ฌ Benchmark Baseline Results

Benchmark results from the accompanying study (Tristansyah & Ichsan, 2026) under deterministic decoding ($\text{Temperature} = 0.0$):

Model ArchitectureScale / RegimeSingle-Turn Tool Acc (%)Single-Turn Strict AST (%)Multi-Turn Strict AST (%)Overall Strict AST (%)Format Validity (%)Latency (s)Hosting Footprint
Claude-Opus-5Frontier / B89.09%87.27%64.42%75.85%99.01%3.94 sCloud API
Gemma4-31BFrontier / A94.55%85.45%64.17%74.81%100.00%1.35 s>28 GB VRAM
GPT-OSS-120BFrontier / A85.45%80.00%59.82%69.91%98.72%1.52 s>75 GB VRAM
GPT-OSS-20BMid-Scale / A89.09%87.27%57.10%72.19%97.44%2.65 s~18 GB VRAM
MiniMax-M3Mid-Scale / B87.27%85.45%21.88%53.67%97.03%5.43 sCloud API
Nemotron-30BMid-Scale / A61.82%61.82%14.24%38.03%68.59%3.58 s>26 GB VRAM
TelkomNusa-LLM-7B (SFT)7B / A89.09%83.64%33.04%58.34%100.00%11.02 s~9.2 GB (NVIDIA L4)
TelkomNusa-SLM-3B (SFT)3B / A85.45%72.73%16.38%44.55%100.00%10.30 s~5.8 GB (NVIDIA T4)
Llama-3.2-3BOpen / B81.82%78.18%3.99%41.09%100.00%22.03 s~6.1 GB (NVIDIA T4)
Qwen2.5-3B-BaseOpen / A67.27%61.82%4.13%32.98%99.01%7.65 s~5.8 GB (NVIDIA T4)
Phi-4-miniOpen / B10.91%10.91%0.00%5.46%97.03%6.68 s~6.8 GB (NVIDIA T4)

Key Insights:

  • โ€”Domain Weight Specialization Outperforms Scale: TelkomNusa-LLM-7B achieves 83.64% Single-Turn AST, surpassing the massive GPT-OSS-120B (80.00%).
  • โ€”Edge Viability: TelkomNusa-SLM-3B achieves 44.55% Overall AST, outperforming Llama-3.2-3B (41.09%) with over $53\%$ lower latency on a commodity NVIDIA Tesla T4 GPU.

๐Ÿ”’ Data Hygiene & Contamination Safeguards

To prevent artificial metric inflation and data leakage:

  1. 1.Strictly Disjoint Testbed: The 757 evaluation scenarios are strictly quarantined (held-out) from the 3,510 training pairs of the TelkomNusa SFT corpus.
  2. 2.Decontamination Protocol: N-gram overlap and AST similarity audits were executed following the contamination protocols established by Deng et al. (NAACL 2024) and Lee et al. (ACL 2022).
  3. 3.Synthetic PII Masking: All MSISDNs utilize synthetic Indonesian telecommunications number allocations (0812xxxx, 0811xxxx). No real customer Personally Identifiable Information (PII) is included, ensuring compliance with Indonesian Law No. 27/2022 (UU PDP).

๐Ÿ“œ How to Cite

If you utilize TelAgentBench-ID in your research, please cite both the accompanying study and the original foundational benchmark:

bibtex
@article{tristansyah2026benchmarking,
  title={Benchmarking Small Language Models Against Frontier Giants for Telecommunications BSS Tool-Calling: Execution Fidelity and Epistemic Calibration},
  author={Tristansyah, M Rislan and Ichsan, Ichwan Nul},
  journal={JUTI: Jurnal Ilmiah Teknologi Informasi},
  volume={x},
  number={x},
  year={2026},
  publisher={Universitas Pendidikan Indonesia \& Institut Teknologi Sepuluh Nopember}
}

@inproceedings{lee2025telagentbench,
  title={TelAgentBench: A Multi-Faceted Benchmark for Evaluating LLM-Based Agents in Telecommunications},
  author={Lee, S. and Potdar, S. and Rojas-Barahona, L. and Montella, S. and others},
  booktitle={Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP): Industry Track},
  pages={1173--1211},
  year={2025},
  publisher={Association for Computational Linguistics},
  doi={10.18653/v1/2025.emnlp-industry.83}
}

๐Ÿ“„ License & Attribution

This dataset is distributed under the Creative Commons Attribution 4.0 International License ([CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)). You are free to share, copy, modify, and build upon this benchmark for academic and commercial purposes, provided appropriate attribution is given to the original authors.