datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
telecom-intent-config-sft-10k
Telecom Intent→Config SFT Dataset (10K)
The first open SFT dataset for training LLMs to translate natural language network intents into structured 5G/6G configurations.
This dataset addresses the #1 gap identified in the telecom LLM research landscape: there is no public training dataset for intent-to-policy translation. All existing telecom datasets (TeleQnA, ORANBench-13K, 6G-Bench) are MCQ evaluation benchmarks — not instruction-following format. This dataset fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/nraptisss/telecom-intent-config-sft-10k.tau2-telecom-agent-sft
τ²-bench telecom — teacher trajectories for agent SFT
784 accepted multi-turn tool-use trajectories on the telecom domain of
τ²-bench, collected to cold-start an 8B model
before reinforcement learning.
Training code, the full lab record and the RL stages that follow are at
yuecao365/tau2telecom_RL.
The point of this domain is dual control: the agent has thirteen backend APIs, the customer
has thirty tools on their own handset, and 76% of the actions a task expects can only be… See the full description on the dataset page: https://huggingface.co/datasets/cy-330/tau2-telecom-agent-sft.
