datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ultrachat-sharegpt-5GBmoe-5g-nrxCascadingCacheArray_702_5GVoai-5g-srs-ranging-dataset
OAI 5G NR SRS Ranging Captures
Uplink Sounding Reference Signal (SRS) channel-estimate captures from a monolithic 5G NR
software-defined-radio testbed, collected for SRS-based ranging experiments.
The gNB (OpenAirInterface on a USRP X410, Band n78, 40 MHz / 106 PRB) configures each UE
to transmit SRS; the gNB's per-SRS frequency-domain channel estimate, oversampled IDFT CIR,
and ToA estimate are streamed off the PHY via OAI's T_tracer and recorded at a series of
known… See the full description on the dataset page: https://huggingface.co/datasets/ahancock516/oai-5g-srs-ranging-dataset.LlaMA3.1-7B-Instruct-64k-finetune-5G-1Oct2024-TEXTLlaMA3.1-7B-Instruct-15000-finetune-5G-1Oct2024-TEXTLlaMA3.1-7B-Instruct-finetune-5G-1Oct2024oai-5g-plkg-dataset
OAI 5G PLKG Paired Dataset
PUSCH (uplink) and PDSCH (downlink) channel-estimate captures from an OAI
5G NR SA testbed (USRP X410 gNB + USRP B210 UE, band n78, 106 PRB, PCI=1),
collected for a channel-informed neural PHY secret-key generation (PLKG)
demo paper. Campaigns 1-3: captures are handshake/RRC-setup-phase only
(strict DMRS-comb filtering rejects connected-mode single-layer data on
both plugins). Campaign 4 relaxes this filtering on both plugins and
includes ordinary… See the full description on the dataset page: https://huggingface.co/datasets/ahancock516/oai-5g-plkg-dataset.LlaMA3.1-7B-Instruct-15k-finetune-5G-10Oct2024-TEXTLlaMA3.1-7B-Instruct-14k-finetune-5G-10Oct2024-TEXTLlaMA3.1-7B-Instruct-13k-finetune-5G-10Oct2024-TEXTqode-smoker-detection5gnr-pusch-iq-dmrs
5G NR PUSCH IQ DMRS Captures
Frequency-domain PUSCH IQ captures from an OAI 5G NR gNB (NI USRP X410 / USRP B210) with per-device IMSI labels embedded directly in each capture record. Collected over-the-air with commercial Quectel RM520N-GL modems and a software-defined USRP B210 UE.
Dataset summary
Filtering: only captures with a strongly visible DMRS RE comb (active/quiet power ratio ≥ 1.5×) are accepted
Format: v4 binary (.bin) — see nr_pusch_capture_oai for… See the full description on the dataset page: https://huggingface.co/datasets/ahancock516/5gnr-pusch-iq-dmrs.raw-advance-driver-monitoringYouTube-Commons-5G-Raw5G_Faults_FullSynthetic dataset of 5G mobile network faults
3gpp-5g-nr-qa
3GPP 5G NR Q&A Dataset
A question-answering dataset derived from 3GPP 38-series technical specifications covering 5G New Radio (NR). Contains ~30,000 Q&A pairs suitable for fine-tuning LLMs on telecommunications domain knowledge.
Dataset Description
Overview
This dataset contains 29,918 instruction-following Q&A pairs generated from 3GPP 38-series specifications. The questions cover technical aspects of 5G New Radio including physical layer procedures, RRC… See the full description on the dataset page: https://huggingface.co/datasets/raoulbia/3gpp-5g-nr-qa.goal-pashto-chat-sharegpt-5GB
📄 goal-pashto-chat-sharegpt-5GB — Pashto ShareGPT‑Style Chat Dataset
A large‑scale, high‑quality Pashto conversational dataset designed for instruction‑tuning, dialogue modeling, and LLM alignment.This dataset contains ~5GB of multi‑turn Pashto conversations inspired by ShareGPT, covering reasoning, advice, education, culture, and general knowledge.
📌 Dataset Summary
goal-pashto-chat-sharegpt-5GB is a curated collection of Pashto user–assistant conversations… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/goal-pashto-chat-sharegpt-5GB.qwen3_sft_5gb
Qwen3-14B SFT — Combined
Merged dataset: original batch (709,027 conversations) + batch 2
(682,614 new conversations, cross-deduplicated against the original
before merging — zero duplicate content between the two).
Batch 2 sources
Capybara, No Robots, Open-Platypus, SlimOrca, Orca-Math, CodeFeedback,
Tulu-3-SFT-Mixture, Dolly-15k
Splits
train: 1,363,808
test: 27,833
Format
messages: [{role, content}] — use… See the full description on the dataset page: https://huggingface.co/datasets/lujain127/qwen3_sft_5gb.raw-vehicle-classification-marquis035Gspeakleash-tokenizer-5gb-sample
SpeakLeash tokenizer 42GB quality sample
Private tokenizer-training sample built from a stratified local SpeakLeash text snapshot on gb10.
Target: 5.0 GiB raw JSONL UTF-8 bytes after document filters and exact text dedup.
Format: data/tokenizer_sample_*.jsonl.zst, one JSON object per line with text, source, source_kind.
This is intended for tokenizer/BPE training convergence tests.
wikipedia_clean_5GB5gdataqode-vehicle-typestelco-5G-data-faultsSynthetic test dataset for 5G data service faults in the core and RAN network domains. It is used to train the telcoLLM to simulate assistance model for network operations.
2_5GBdataNeversNet5G_Vehicular_5G_NR_Dataset
NeversNet5G
NeversNet5G is a city-scale 5G NR vehicular network dataset generated through
SUMO and Simu5G/OMNeT++ co-simulation over the real urban road network of
Nevers, France. The release contains processed per-UE event-level CSV files,
not the raw OMNeT++ .vec/.sca outputs.
Release Contents
data/: 716 per-UE metric CSV files, grouped by scenario part.
configs/: OMNeT++/Simu5G configuration files used for the scenario. The
released omnetpp.ini is trimmed to… See the full description on the dataset page: https://huggingface.co/datasets/Askedrin/NeversNet5G_Vehicular_5G_NR_Dataset.Wikipedia_5gram_less_orders
Dataset Card for "Wikipedia_5gram_less_orders"
More Information needed
CulturaY_vi_5GB
