Rislantrs/TelAgentBench-ID
TelAgentBench-ID: A Comprehensive Benchmark for Evaluating Autonomous LLM Agents in Telecommunications Business Support Systems ๐ Dataset Summary TelAgentBench-ID is the first comprehensive, multi-faceted benchmark specifically constructed to evaluate the Action Execution Fidelity and Epistemic Calibration of Large Language Models (LLMs) and Small Language Models (SLMs) within the Telecommunications Business Support Systems (BSS) domain in Indonesian.โฆ See the full description on the dataset page: https://huggingface.co/datasets/Rislantrs/TelAgentBench-ID.
Update dataset configs to prioritize BSS tool-calling actions
Delete validation.json with huggingface_hub
Release TelAgentBench-ID: Comprehensive Indonesian Telecom BSS Benchmark
Release TelAgentBench-ID: Comprehensive Indonesian Telecom BSS Benchmark
Release TelAgentBench-ID: Comprehensive Indonesian Telecom BSS Benchmark
Release TelAgentBench-ID: Comprehensive Indonesian Telecom BSS Benchmark
initial commit
