datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
netopsbench-trace
NetOpsBench Agent Traces
This dataset contains sanitized NetOpsBench benchmark trace artifacts.
The legacy cross-model snapshot contains minimal-deepagent runs for
MiniMax M3, DeepSeek V4 Pro, Kimi K2.6, and OpenAI GPT-5.5 on the XS, Small,
Medium, and Large CLOS profiles.
The NetOpsBench v0.2 release adds a separately versioned
minimal-deepagent / deepseek-v4-pro snapshot across all seven built-in
profiles: XS, Small, Medium, Large, Xlarge, Fat-tree K=8, and Fat-tree K=12.
It… See the full description on the dataset page: https://huggingface.co/datasets/yyyyyt/netopsbench-trace.NetSecDataDetails about the creation of this dataset can be found in the article Hackphyr: A Local Fine-Tuned LLM Agent for Network Security Environments.
NetBench
NetBench Dataset
Dataset Overview
The NetBench Dataset is a curated collection of expert-level question-answer pairs designed to benchmark the ability of large language models (LLMs) to achieve network subject matter expert (SME) intelligence across 20 critical telecommunications and network engineering categories. These categories include:
Network Fundamentals & L2 Switching: Basic device access, Layer 2 concepts (VLANs, STP, LAG), L2 security, and interface… See the full description on the dataset page: https://huggingface.co/datasets/NetoAISolutions/NetBench.Amawal.net-Dataset
Dataset Card for Amawal.net-Dataset
This dataset compiles the crowdsourced lexicon and linguistic entries from the legacy platform Amawal.net. Following the closure of the original website, this repository provides an archive of the community's multi-dialectal contributions to preserve its linguistic value and make it accessible for natural language processing (NLP), lexicography, and digital humanities applications.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amawal.net-Dataset.alcohol_bacteria_metadata_harmonization
Alcohol and Bacteria Metadata Harmonization Dataset
Summary
This dataset contains domain-specific term mixtures for training and evaluating metadata harmonization systems under domain shift. Each configuration includes a defined ratio of alcohol-related and bacteria-related terms to support experiments on generalization and domain adaptation. Each entry includes a term representation, its corresponding harmonized standard, and metadata such as variation type and source… See the full description on the dataset page: https://huggingface.co/datasets/netrias/alcohol_bacteria_metadata_harmonization.code_search_net_kotlin
Dataset Summary
This dataset was converted based on code_search_net (https://huggingface.co/datasets/code-search-net/code_search_net)
Languages
Kotlin programming language
C++ programming language
Data Instances
A data point consists of a function code along with its documentation. Each data point also contains meta data on the function, such as the repository it was extracted from.
{
'id': '0',
'repository_name': 'organisation/repository'… See the full description on the dataset page: https://huggingface.co/datasets/namyi/code_search_net_kotlin.cancer_metadata_harmonization
Cancer Metadata Harmonization Dataset
Summary
This dataset contains cancer-related terms for training and evaluating metadata harmonization systems in the biomedical domain. Each entry includes a term representation, its corresponding harmonized standard, and metadata such as semantic type, variation type, and source terminology. Term representations include standard forms as well as lexical variations (e.g., synonyms, abbreviations) and are harmonized to biomedical… See the full description on the dataset page: https://huggingface.co/datasets/netrias/cancer_metadata_harmonization.Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_TextsHere presented a partially synthesized dataset, developed utilizing the GPT-4 model, for the purpose of NLG, particulary for the task of hierarchical generation of longer texts from short summaries. The creation of this dataset was undertaken as a component of my thesis paper. It incorporates excerpts from prominent British and American novels, from which plots, summaries, and metadata have been derived using GPT-4 API to facilitate extensive future research.
The metadata included in the… See the full description on the dataset page: https://huggingface.co/datasets/Fleur-roar/Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_Texts.
