cy-330/tau2-telecom-agent-sft
τ²-bench telecom — teacher trajectories for agent SFT 784 accepted multi-turn tool-use trajectories on the telecom domain of τ²-bench, collected to cold-start an 8B model before reinforcement learning. Training code, the full lab record and the RL stages that follow are at yuecao365/tau2telecom_RL. The point of this domain is dual control: the agent has thirteen backend APIs, the customer has thirty tools on their own handset, and 76% of the actions a task expects can only be… See the full description on the dataset page: https://huggingface.co/datasets/cy-330/tau2-telecom-agent-sft.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face