CoolFace
Datasetpublic

GSMA/Telco-Common-Corpus

Telco Common Corpus (TCC) is a ten billion tokens collection of fully open, free licensed telecommunications knowledge (scientific literature, patents, open data, and open-web projects) with licence and provenance verified at a document-level. TCC stems from GSMA's effort to make AI work for the telecom sector. The Open-Telco LLM Benchmarks and the broader Open Telco AI initiative have already established that current models fall short on real telecom tasks, including network management and… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/Telco-Common-Corpus.

sourceHugging Faceupdated 3mo agoView on Hugging Face
4likes763downloads
7 commits on main
c590e4e3mo ago

Update README.md

Pclanglais
3c3975a3mo ago

Delete tcc_sample.parquet

Pclanglais
2d247b63mo ago

Initial upload: 100 sharded parquets (1.78M docs, ~9.7B tokens)

Pclanglais
25a59b53mo ago

Update README.md

Pclanglais
8fa13343mo ago

Upload tcc_sample.parquet

Pclanglais
d9bc2b83mo ago

Create README.md

Pclanglais
1d223cd3mo ago

initial commit

Pclanglais