CoolFace
Datasetpublic

Rislantrs/TelAgentBench-ID

TelAgentBench-ID: A Comprehensive Benchmark for Evaluating Autonomous LLM Agents in Telecommunications Business Support Systems ๐Ÿ“Œ Dataset Summary TelAgentBench-ID is the first comprehensive, multi-faceted benchmark specifically constructed to evaluate the Action Execution Fidelity and Epistemic Calibration of Large Language Models (LLMs) and Small Language Models (SLMs) within the Telecommunications Business Support Systems (BSS) domain in Indonesian.โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelAgentBench-ID.

sourceHugging Facecc-by-4.0updated 6d agoView on Hugging Face
0likes579downloads
7 commits on main
ecfbb3d6d ago

Update dataset configs to prioritize BSS tool-calling actions

Rislantrs
5e7eac06d ago

Delete validation.json with huggingface_hub

Rislantrs
8b5c14f6d ago

Release TelAgentBench-ID: Comprehensive Indonesian Telecom BSS Benchmark

Rislantrs
584321b6d ago

Release TelAgentBench-ID: Comprehensive Indonesian Telecom BSS Benchmark

Rislantrs
c8714b16d ago

Release TelAgentBench-ID: Comprehensive Indonesian Telecom BSS Benchmark

Rislantrs
785772a6d ago

Release TelAgentBench-ID: Comprehensive Indonesian Telecom BSS Benchmark

Rislantrs
0ef5df16d ago

initial commit

Rislantrs