datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
telco-gaia
Telco-GAIA
A GAIA-style benchmark for AI agents operating over a real telecom operator's
website snapshot plus a synthetic customer database. 100 tasks across 7
categories: Pricing, Miscellaneous, Images, Web Archives, PDF, PDF Visual,
Database.
Agents read questions.json + environment.md, browse the local website
(:8080) and query the database API (:8081), and produce a GAIA-compatible
submission.json.
What's here
File
What… See the full description on the dataset page: https://huggingface.co/datasets/kaust-generative-ai/telco-gaia.Telco-Troubleshooting-Agentic-Challenge
Telco Troubleshooting and Optimization Agentic Challenge
[!NOTE]
IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing the dataset.
Telco Troubleshooting Agentic Challenge focuses on network operations and maintenance. Participants are required to build intelligent agents that complete network fault diagnosis and troubleshooting tasks in the contest of wireless and IP networks. More specifically, the goal will… See the full description on the dataset page: https://huggingface.co/datasets/netop/Telco-Troubleshooting-Agentic-Challenge.open_telco_v2_customopen-telco-energy-logs
Open-Telco Energy Logs
Per-model GPU energy measurements for 67 open-weight LLMs evaluated on the Open Telco benchmark suite. This is the raw data backing the energy and AMEI (AI Model Energy Index) tables in the AMEI paper.
Three measurement campaigns are included:
Campaign
Models
Hardware
Runtime
Suite
Samples/model
Schema
single_h100/
24
1×H100-80GB SXM
vLLM 0.8.5 (TP=1)
open_telco_5_benchmarks
1,400
v1.0
4x_h100/
35
4×H100-80GB SXM
vLLM 0.15.1 (TP=4)… See the full description on the dataset page: https://huggingface.co/datasets/emolero/open-telco-energy-logs.telco-5G-data-faultsSynthetic test dataset for 5G data service faults in the core and RAN network domains. It is used to train the telcoLLM to simulate assistance model for network operations.
gsma-sft-stage1-v4-datasetOpen-Telco-1telco-troubleshooting-tracestelco-5G-data-faultsSynthetic test dataset for 5G data service faults in the core and RAN network domains. It is used to train the telcoLLM to simulate assistance model for network operations.
telco-5g-QnA
Telco 5G Data set
The data set consists of 125 questions specific to 5G, OSS/BSS. With categories of Enterprise, 5G, OSS etc.
telco-gaia-groundtruth
Telco-GAIA — Ground Truth (gated)
Answers + gold reasoning steps for the 100 Telco-GAIA tasks, plus the scorer.
ground_truth.json — task_id, category, language, question, final_answer, answer_type, steps, tools
evaluate.py, answer_matching.py — GAIA-style exact-match scorer
Pair with the open dataset kaust-generative-ai/telco-gaia for the questions, website,
and harness.
python evaluate.py --submission submission.json --ground-truth ground_truth.json --output results.json
Training_TAO_V2_SFTopen_telco_customtelco-improvedrhfl-telcotelcom_services_sitentico
empgces/telcom_services_sitentico
Dataset de SFT para tarifários de voz. Cada exemplo inclui a FONTE DE VERDADE (JSON do plano) embutida na prompt.
Formato
unitel_sft_prompt_completion.jsonl(.gz): prompt, completion
unitel_sft_chat_with_source.jsonl(.gz): chat com a mesma prompt no user.
Campos esperados nas Q&A
tariff_id, question, answer, opcional persona.
