datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TelecomTS
📡 TelecomTS: A Multi-Modal Telecom Dataset
TelecomTS is a large-scale, high-resolution, multi-modal dataset derived from a 5G telecommunications testbed. It is the first public observability dataset to preserve deanonymized observability metrics with absolute scale information, encompassing by design various downstream tasks beyond forecasting such as anomaly detection, root-cause analysis, and multi-modal reasoning.
Observability data, particularly in… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/TelecomTS.creative_writing
Dataset Card for telecomadm1145/creative_writing
Dataset Details
Dataset Description
This dataset is a small-scale instruction–response dataset focused on creative writing tasks.Each example consists of a prompt (instruction specifying writing style, perspective, tone, etc.) and a response (a story segment or novel-like output).
The dataset emphasizes:
Creative Writing (light novel style, emotional narrative, dialogue-driven, descriptive prose).… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/creative_writing.telecom
telecom
An executable Environment for evaluating and training tool-using agents. Kullback built it from recorded traces of a working agent, and Leibler publishes it.
The package holds the rebuilt world: a database, one function per tool that behaves the way the real tool was observed to behave, the compiled policy, and the Starting state each Task begins from. It also holds the Task list with the instruction a candidate gets, and a code-only Verifier per Task that grades the… See the full description on the dataset page: https://huggingface.co/datasets/leibler/telecom.tau2-telecom-agent-sft
τ²-bench telecom — teacher trajectories for agent SFT
784 accepted multi-turn tool-use trajectories on the telecom domain of
τ²-bench, collected to cold-start an 8B model
before reinforcement learning.
Training code, the full lab record and the RL stages that follow are at
yuecao365/tau2telecom_RL.
The point of this domain is dual control: the agent has thirteen backend APIs, the customer
has thirty tools on their own handset, and 76% of the actions a task expects can only be… See the full description on the dataset page: https://huggingface.co/datasets/cy-330/tau2-telecom-agent-sft.TelecomTS
📡 TelecomTS: A Multi-Modal Telecom Dataset
TelecomTS is a large-scale, high-resolution, multi-modal dataset derived from a 5G telecommunications testbed. It is the first public observability dataset to preserve deanonymized observability metrics with absolute scale information, encompassing by design various downstream tasks beyond forecasting such as anomaly detection, root-cause analysis, and multi-modal reasoning.
Observability data, particularly in… See the full description on the dataset page: https://huggingface.co/datasets/Govisha/TelecomTS.africa-synth-mobile-telecom-mobile-subscriber-data-all
African Mobile Subscriber Data | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: json - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-mobile-telecom-mobile-subscriber-data-all.esjzone_2024
Dataset Card for telecomadm1145/esjzone_2024
Dataset Details
Dataset Description
Novels from Esjzone.
Check telecomadm1145/creative_writing for instruction finetuning and telecomadm1145/esjzone_2024_chunked_8k for continuation.
Curated by: telecomadm1145
Language(s) (NLP): Chinese
License: MIT
Dataset Sources
Source Data: Publicly available online novels (especially light novels).
Uses
Better using DPP similar to below code:… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/esjzone_2024.TelecomTS
📡 TelecomTS: A Multi-Modal Telecom Dataset
TelecomTS is a large-scale, high-resolution, multi-modal dataset derived from a 5G telecommunications testbed. It is the first public observability dataset to preserve deanonymized observability metrics with absolute scale information, encompassing by design various downstream tasks beyond forecasting such as anomaly detection, root-cause analysis, and multi-modal reasoning.
Observability data, particularly in telecommunications, differs… See the full description on the dataset page: https://huggingface.co/datasets/SamalaSharan/TelecomTS.test0
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/test0.telecom-customer-support-synthetic-replicas
Customer Support Differentially Private Synthetic Conversations Dataset
This dataset contains pairs of customer support conversations: original conversations and their synthetic counterparts generated with differential privacy (DP) guarantees (ε=8.0, δ=1e-5). The conversations cover technical support topics related to performance and speed concerns in mobile devices.
Dataset Description
Overview
The Customer Support Differentially Private Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Ming-secludy/telecom-customer-support-synthetic-replicas.telecom-rag-faiss-last-V2telecom-rag-embeddings-bge-en-icltalkmap-telecom-tajik
Saidzoda Lab — Gated Research Dataset
Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz).
Dataset contents, provenance, and statistics are not publicly disclosed.
Access is granted manually on request.
telecom-rag-faisstelecom-rag-faiss-lastnextgen-telecom-incidentstelecomregulatortelecom-rag-embeddings-bge-base-en
