datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
creative_writing
Dataset Card for telecomadm1145/creative_writing
Dataset Details
Dataset Description
This dataset is a small-scale instruction–response dataset focused on creative writing tasks.Each example consists of a prompt (instruction specifying writing style, perspective, tone, etc.) and a response (a story segment or novel-like output).
The dataset emphasizes:
Creative Writing (light novel style, emotional narrative, dialogue-driven, descriptive prose).… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/creative_writing.telecom-conversation-corpus
Telecom 200k Dataset Overview
This dataset consists of 200,000 synthetically generated conversations in a customer service setting for the telecom industry. There are two speakers: a customer, and an agent.
telecom-intent-config-sft-10k
Telecom Intent→Config SFT Dataset (10K)
The first open SFT dataset for training LLMs to translate natural language network intents into structured 5G/6G configurations.
This dataset addresses the #1 gap identified in the telecom LLM research landscape: there is no public training dataset for intent-to-policy translation. All existing telecom datasets (TeleQnA, ORANBench-13K, 6G-Bench) are MCQ evaluation benchmarks — not instruction-following format. This dataset fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/nraptisss/telecom-intent-config-sft-10k.tau2-telecom-agent-sft
τ²-bench telecom — teacher trajectories for agent SFT
784 accepted multi-turn tool-use trajectories on the telecom domain of
τ²-bench, collected to cold-start an 8B model
before reinforcement learning.
Training code, the full lab record and the RL stages that follow are at
yuecao365/tau2telecom_RL.
The point of this domain is dual control: the agent has thirteen backend APIs, the customer
has thirty tools on their own handset, and 76% of the actions a task expects can only be… See the full description on the dataset page: https://huggingface.co/datasets/cy-330/tau2-telecom-agent-sft.esjzone_2024
Dataset Card for telecomadm1145/esjzone_2024
Dataset Details
Dataset Description
Novels from Esjzone.
Check telecomadm1145/creative_writing for instruction finetuning and telecomadm1145/esjzone_2024_chunked_8k for continuation.
Curated by: telecomadm1145
Language(s) (NLP): Chinese
License: MIT
Dataset Sources
Source Data: Publicly available online novels (especially light novels).
Uses
Better using DPP similar to below code:… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/esjzone_2024.telecom-conversation-corpustelecom-customer-support-synthetic-replicas
Customer Support Differentially Private Synthetic Conversations Dataset
This dataset contains pairs of customer support conversations: original conversations and their synthetic counterparts generated with differential privacy (DP) guarantees (ε=8.0, δ=1e-5). The conversations cover technical support topics related to performance and speed concerns in mobile devices.
Dataset Description
Overview
The Customer Support Differentially Private Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Ming-secludy/telecom-customer-support-synthetic-replicas.telecom-conversation-corpus
Telecom 200k Dataset Overview
This dataset consists of 200,000 synthetically generated conversations in a customer service setting for the telecom industry. There are two speakers: a customer, and an agent.
