datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ub-networking-dataset-2024-2PRCA-Net-dataset
Ray-Traced Cross-Frequency Radio Map Dataset
A large ray-traced radio-map (path-loss) dataset for zero-shot
cross-frequency generalization research, generated with
Sionna RT over real urban geometry from
OpenStreetMap.
150 urban scenes across 15 cities, 256×256 rasters
8 transmitters per scene across three deployment strata (street,
rooftop, mast)
6 carrier frequencies: 1.8, 3.5, 7, 28 GHz (training) + 10, 60 GHz
(held out, for interpolation / extrapolation studies)
7,200… See the full description on the dataset page: https://huggingface.co/datasets/SHussain37/PRCA-Net-dataset.CrediBench
Dataset Card for CrediBench 1.1
CrediBench is a large-scale, temporal webgraph constituted of web data pulled from Common Crawl.
Dataset Details
Dataset Description
This dataset is composed of monthly slices of large-scale web networks. These webgraphs contain 1+ billion edges, and 45+ million nodes per month.
In these webgraphs, the nodes represent a website domain (e.g, google.com) and an edge represents a directed hyperlink relation (e.g, an… See the full description on the dataset page: https://huggingface.co/datasets/credi-net/CrediBench.Network-Intrusion-Detection-DataTON_IoT_network
TON IoT Network
The TON IoT train test network dataset provided by https://research.unsw.edu.au/projects/toniot-datasets
Dataset Details
The datasets have been called 'ToN_IoT' as they include heterogeneous data sources collected from Telemetry datasets of IoT and IIoT sensors, Operating systems datasets of Windows 7 and 10 as well as Ubuntu 14 and 18 TLS and Network traffic datasets. The datasets were collected from a realistic and large-scale network designed at the… See the full description on the dataset page: https://huggingface.co/datasets/codymlewis/TON_IoT_network.code_search_net_python_10000_examplesTV-Shows-Netflix-Disney
Dataset: Series y Películas de Netflix y Disney+
Este dataset contiene información sobre series disponibles en las plataformas Netflix y Disney+. Está diseñado para análisis simples y exploratorios, ya que cuenta con 12 columnas que describen las características básicas de cada serie.
Estructura del Dataset
El dataset incluye las siguientes columnas:
show_id: Identificador único de cada serie.
type: Tipo de contenido (en este caso, "TV Show").
title: Nombre de la serie p… See the full description on the dataset page: https://huggingface.co/datasets/MarcoM003/TV-Shows-Netflix-Disney.netflix-shows
Dataset Card for Dataset: NetFlix Shows
Dataset Summary
The raw data is Web Scrapped through Selenium. It contains Unlabelled text data of around 9000 Netflix Shows and Movies along with Full details like Cast, Release Year, Rating, Description, etc.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/netflix-shows.NFI_FARED_IMUThis is the README file for the dataset Netherlands Forensic Institute: Forensic Activity Recognition Dataset (NFI_FARED), published as a part of the paper "Hi-OSCAR: Hierarchical Open-set Classifier for Human Activity Recognition.". Two forms of data were collected: Digital Traces from iPhones worn on the subjects' bodies, and raw sensor signals from body-worn Inertial Measurement Units (IMUs). This dataset and README refers to the IMU data. The Digital Trace data is available here.
NFI_FARED… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/NFI_FARED_IMU.network-packet-flow-header-payloadEach row contains the information of a network packet and its label. The format is given below:
NADW-network-attacks-dataset
Network Traffic Dataset for Anomaly Detection
Overview
This project presents a comprehensive network traffic dataset used for training AI models for anomaly detection in cybersecurity. The dataset was collected using Wireshark and includes both normal network traffic and various types of simulated network attacks. These attacks cover a wide range of common cybersecurity threats, providing an ideal resource for training systems to detect and respond to real-time network… See the full description on the dataset page: https://huggingface.co/datasets/onurkya7/NADW-network-attacks-dataset.netopsbench-trace
NetOpsBench Agent Traces
This dataset contains sanitized NetOpsBench benchmark trace artifacts.
The legacy cross-model snapshot contains minimal-deepagent runs for
MiniMax M3, DeepSeek V4 Pro, Kimi K2.6, and OpenAI GPT-5.5 on the XS, Small,
Medium, and Large CLOS profiles.
The NetOpsBench v0.2 release adds a separately versioned
minimal-deepagent / deepseek-v4-pro snapshot across all seven built-in
profiles: XS, Small, Medium, Large, Xlarge, Fat-tree K=8, and Fat-tree K=12.
It… See the full description on the dataset page: https://huggingface.co/datasets/yyyyyt/netopsbench-trace.Marvel_network
Dataset Card for Marvel Network
This is a dataset for Marvel universe social network, which contains the relationships between Marvel heroes.
Dataset Description
The Marvel Comics character collaboration graph was originally constructed by Cesc Rosselló, Ricardo Alberich, and Joe Miro from the University of the Balearic Islands. They compare the characteristics of this universe to real-world collaboration networks, such as the Hollywood network, or the one created by… See the full description on the dataset page: https://huggingface.co/datasets/ShimizuYuki/Marvel_network.proto-social-network-canal-barra
Canal Barra Digital Archaeology Dataset
This dataset preserves structured historical evidence related to Canal Barra, a Brazilian digital community founded in 1996 around the #barra IRC channel on the BRASnet network.
Canal Barra combined IRC communication, web-based profiles, persistent nicknames, access-level governance and recurring in-person meetings in Rio de Janeiro. The dataset supports historical and academic investigation into Canal Barra as an early… See the full description on the dataset page: https://huggingface.co/datasets/raphaelnercessian/proto-social-network-canal-barra.netzero_reduction_dataNetSecDataDetails about the creation of this dataset can be found in the article Hackphyr: A Local Fine-Tuned LLM Agent for Network Security Environments.
netflix-medianetwork-dns-resolution-coherence-risk-v0.1What this repo is for
Detect DNS instability before services fail.
Covers real operational signals:
rising resolution latency
SERVFAIL spikes
authoritative mismatch
cache poisoning
missing failover resolvers
Used by:
ISPs
cloud providers
enterprises
SRE teams
benchmark-dataset-different-gpu-workload
GPU catalog × LLM workload VRAM benchmark
Summary
Tabular benchmark in CSV form: each row pairs a catalog GPU (gpu_id, gpu_display_name, catalog_gpu_vram_gb) with a concrete LLM inference-style workload (model, parameter count, context length, precision, batch size, concurrent users). The file records math_engine VRAM component estimates (weights, KV cache, activations, overhead, totals, tier), a document_engine recommended VRAM value, a short comparison summary… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-dataset-different-gpu-workload.NetFlix-Imdb-Engagements-FilmsNetBench
NetBench Dataset
Dataset Overview
The NetBench Dataset is a curated collection of expert-level question-answer pairs designed to benchmark the ability of large language models (LLMs) to achieve network subject matter expert (SME) intelligence across 20 critical telecommunications and network engineering categories. These categories include:
Network Fundamentals & L2 Switching: Basic device access, Layer 2 concepts (VLANs, STP, LAG), L2 security, and interface… See the full description on the dataset page: https://huggingface.co/datasets/NetoAISolutions/NetBench.network-security-route-hijack-coherence-risk-v0.1What this repo is for
Detect routing security incidents fast.
Covers:
unexpected origin ASN
invalid ROAs
AS-path anomalies
suspicious more-specific prefixes
observed traffic diversion
whether mitigation happened
This is high-impact because one leak can break many networks.
Amawal.net-Dataset
Dataset Card for Amawal.net-Dataset
This dataset compiles the crowdsourced lexicon and linguistic entries from the legacy platform Amawal.net. Following the closure of the original website, this repository provides an archive of the community's multi-dialectal contributions to preserve its linguistic value and make it accessible for natural language processing (NLP), lexicography, and digital humanities applications.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amawal.net-Dataset.network-config-intent-state-coherence-drift-v0.1What this repo is for
Detect when you think you deployed a change but the network did not converge to the intended state.
Daily failure modes:
partial deployment
stale ACLs or route-maps
“push succeeded” but running config differs
rollback does not restore baseline
inconsistent policy across devices
This dataset trains a model to flag drift fast.
alpaca_dataset_i_foundzero_trust_network_security_logsaviation-network-delay-cascade-coherence-risk-v0.1What this repo is for
Detect when a delay at one airport
spreads across the network.
Flags
low buffer with rising hub delay
crew or aircraft rotation tight
downstream delay spike
cancellation cascade
Foreign-Direct-Investment-Net-Inflows-Current-USD-Africa
Foreign Direct Investment Net Inflows Current USD Africa | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Foreign-Direct-Investment-Net-Inflows-Current-USD-Africa.benchmark-finetune-dpo-v1
Odyn benchmark: DPO LoRA fine-tuning peak VRAM (V1)
Curated benchmark rows for validating GPU memory estimators during DPO + LoRA fine-tuning. Each row pairs a published or measured expected peak VRAM with inputs to a math engine (model size, context length, batch, LoRA rank, precision, parallelism) plus optional VRAM breakdown and provenance.
This dataset is not preference-pair training JSONL (UltraFeedback-style). It is evaluation ground truth for placement / scheduler memory… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-finetune-dpo-v1.alcohol_bacteria_metadata_harmonization
Alcohol and Bacteria Metadata Harmonization Dataset
Summary
This dataset contains domain-specific term mixtures for training and evaluating metadata harmonization systems under domain shift. Each configuration includes a defined ratio of alcohol-related and bacteria-related terms to support experiments on generalization and domain adaptation. Each entry includes a term representation, its corresponding harmonized standard, and metadata such as variation type and source… See the full description on the dataset page: https://huggingface.co/datasets/netrias/alcohol_bacteria_metadata_harmonization.
