datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_search_net
Dataset Card for CodeSearchNet corpus
Dataset Summary
CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages.
CodeSearchNet corpus was gathered to support the CodeSearchNet challenge, to explore the problem of code retrieval using natural language.
Supported Tasks and Leaderboards
language-modeling: The dataset can be used to… See the full description on the dataset page: https://huggingface.co/datasets/code-search-net/code_search_net.llm-network-study-data
LLM-Network-Study-Data
Per-request network captures (.pcapng) collected by the
LLM-Network-Study benchmark harness (benchmark.py and the
per-workload test scripts). Each directory holds one capture file per request,
named request_<id>_run<n>_<timestamp>.pcapng.
A directory name encodes four dimensions:
<capture-env>_<provider/model>_<workload>[_<dataset/variant>]_results
Dimension legend
Dimension
Values
Meaning
Capture env
ethernet
Wired connection to… See the full description on the dataset page: https://huggingface.co/datasets/wayslab/llm-network-study-data.Vera-Layered-Video-Dataset
Dataset for Vera: A Layered Diffusion Model for Content-Preserving Video Editing
Hongkai Zheng¹²* ·
Ta-Ying Cheng² ·
Benjamin Klein² ·
Yisong Yue¹ ·
Zhuoning Yuan²†
¹California Institute of Technology ²Netflix, Inc.
*Work done during an internship at Netflix †Project Lead
TL;DR: A layered diffusion framework for video editing. Vera jointly generates an edit layer, an alpha… See the full description on the dataset page: https://huggingface.co/datasets/netflix/Vera-Layered-Video-Dataset.NetEaseCrowd
🧑🤝🧑 NetEaseCrowd: A Dataset for Long-term and Online Crowdsourcing Truth Inference
View it in GitHub
Introduction
We introduce NetEaseCrowd, a large-scale crowdsourcing annotation dataset based on
a mature Chinese data crowdsourcing platform of NetEase Inc..
NetEaseCrowd dataset contains about 2,400 workers, 1,000,000 tasks, and 6,000,000 annotations between them,
where the annotations are collected in about 6 months.
In this dataset, we provide ground truths for… See the full description on the dataset page: https://huggingface.co/datasets/liuhyuu/NetEaseCrowd.NetSpec-LLM
📁 Network Spec for LLM Understanding
📄 Overview
This repository houses a comprehensive collection of ETSI (European Telecommunications Standards Institute) documents, systematically downloaded, processed, and organized for streamlined access and analysis. Each ETSI deliverable is paired with its corresponding metadata to ensure thorough information management.
🔍 Data Processing Workflow
The data processing involves two main scripts that automate the… See the full description on the dataset page: https://huggingface.co/datasets/rasoul-nikbakht/NetSpec-LLM.code-search-net-python
Dataset Card for "code-search-net-python"
Dataset Description
Homepage: None
Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python
Paper: None
Leaderboard: None
Point of Contact: @Nan-Do
Dataset Summary
This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-python.knee_fast_mri
Dataset for AVS-Net Pre-training
The dataset utilized in the pre-training of the AVS-Net: Attention-based Variable Splitting Network for P-MRI Acceleration model, developed by Y Zhang, J Li, Z Wang, J Duan, and J Li, incorporates data from five distinct protocol sequences. These are:
(coronal_pd)Coronal Spin Density-weighted without Fat Suppression
(coronal_pd_fs)Coronal Spin Density-weighted with Fat Suppression
(sagittal_pd)Sagittal Spin Density-weighted
(sagittal_t2)Sagittal… See the full description on the dataset page: https://huggingface.co/datasets/AVS-Net/knee_fast_mri.actuator_netGUI-Net-1M-extendedGUI-Net-1M
Check more details at how to use this dataset at our repo
GUI-Net-1M is the dataset we keep running the pipeline introduced from TongUI paper.
Due to large file size, we have to split image files into parts. To do the extraction of images, please use the following script:
#!/bin/bash
# Directory containing the split files
SPLIT_DIR="/mnt/bofeidisk2/tmp/baidu_experience_full/images/split_parts_baidu_experience"
OUTPUT_DIR="merged_files"
# Create output directory if it doesn't… See the full description on the dataset page: https://huggingface.co/datasets/Bofeee5675/GUI-Net-1M.bitcoin-network-propagation
Bitcoin network propagation
Timestamped block and transaction announcements received from connected Bitcoin peers, with peer metadata and advertised relay-fee floors. These observations support analysis of announcement timing and differences between connected peers.
Contents
Table
Record
bitcoin_block_announcements
A peer's announcement of a block, timestamped on receipt
bitcoin_transaction_announcements
A retained transaction announcement from a peer… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/bitcoin-network-propagation.vuurwerkverkenner-development-data
NFI Fireworks Development Dataset for the "Vuurwerkverkenner" Application
The Netherlands Forensic Institute (NFI) Fireworks development dataset consists of scans of fireworks wrappers from
fireworks that were investigated in casework in the Netherlands from 2010 onwards. Artificially created snippets
are available for all wrappers, and for a subset of the wrappers photographs of actual fireworks snippets (pieces of the
wrapper post-detonation) are included.
Data… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/vuurwerkverkenner-development-data.bitcoin-cash-network-propagation
Bitcoin Cash network propagation
Timestamped block and transaction announcements received from connected Bitcoin Cash peers, with peer metadata and advertised relay-fee floors. These observations support analysis of announcement timing and differences between connected peers.
Contents
Table
Record
bitcoin_cash_block_announcements
A peer's announcement of a block, timestamped on receipt
bitcoin_cash_transaction_announcements
A retained transaction… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/bitcoin-cash-network-propagation.dogecoin-network-propagation
Dogecoin network propagation
Timestamped block and transaction announcements received from connected Dogecoin peers, with peer metadata and advertised relay-fee floors. These observations support analysis of announcement timing and differences between connected peers.
Contents
Table
Record
dogecoin_block_announcements
A peer's announcement of a block, timestamped on receipt
dogecoin_transaction_announcements
A retained transaction announcement from a… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/dogecoin-network-propagation.DNANet_2p5pMixture_PPF6C_2024
2p5p Mixture DNA Research dataset
This dataset repository contains RFU signal reading (.hid) files and their corresponding person mixture labels (.txt) used in the DNANet paper and code.
The data consist of DNA sample mixtures of 2 to 5 persons, of which the mixtures composition is known,
allowing for training on actual ground-truth data for DNA annotation tools such as DNANet
If you use this dataset in your research please cite it appropriately:
@ARTICLE{Benschop2019,
title… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/DNANet_2p5pMixture_PPF6C_2024.GUI-Net-Mini
TongUI-143K
Training dataset for TongUI: Building Generalized GUI Agents by Learning from Multimodal Web Tutorials
Dataset
Number
TongUI Collected
143 K
Other
237 K
Dataset Introduction
The datasets contains two types of files:
*.json files which is the instructional following data for GUI Task.
*.zip.part file which are GUI screenshots.
For ease of training, this *.json files follow the dataset settings of LLaMA-Factory.
There are two types of GUI… See the full description on the dataset page: https://huggingface.co/datasets/Bofeee5675/GUI-Net-Mini.litecoin-network-propagation
Litecoin network propagation
Timestamped block and transaction announcements received from connected Litecoin peers, with peer metadata and advertised relay-fee floors. These observations support analysis of announcement timing and differences between connected peers.
Contents
Table
Record
litecoin_block_announcements
A peer's announcement of a block, timestamped on receipt
litecoin_transaction_announcements
A retained transaction announcement from a… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/litecoin-network-propagation.3GPP-R18rhizomorphic-networks-data
Rhizomorphic Networks — Data and Analysis Outputs
This dataset contains the experimental image data, segmentation outputs, and
downstream analysis results used for the quantitative analysis of
Armillaria gallica rhizomorphic networks.
The directory structure is organised according to the main stages of the analysis
pipeline:
Rhizomorphic Networks/
│
├── 01_raw inputs/
│ ├── control/
│ ├── furnace/
│ └── nutrient density/
│
├── 02_segmentation outputs/
│ ├── raw… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/rhizomorphic-networks-data.network_security_questionsThis dataset contains a single file full of network security questions in Chinese.
Could be used as good initial sources for scrapers, though not good as your browsing history.
NetConfEval
NetConfEval: Can LLMs Facilitate Network Configuration?
What is it?
We present a set of benchmarks (NetConfEval) to examine the effectiveness of different models in facilitating and automating network configuration described in our paper "NetConfEval: Can LLMs Facilitate Network Configuration?".
📜 Paper - GitHub Repository
This repository contains pre-generated datasets for each of the benchmark task, so that they can be used independently from our testing environment.… See the full description on the dataset page: https://huggingface.co/datasets/NetConfEval/NetConfEval.bridge_network_open_clipNetworking_Commands_DatasetNetworking Commands Dataset
Overview
This dataset is (networking_dataset) contains 750 unique Cisco-specific and general networking commands (NET001–NET750), designed for red teaming AI models in cybersecurity. It focuses on testing model understanding, detecting malicious intent, and ensuring safe responses in enterprise networking environments. The dataset includes both common and obscure commands, emphasizing advanced configurations for adversarial testing.
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Networking_Commands_Dataset.CDB_DEC2024-CochranSampled_Gemma-300m_Embrobo_net_rawresume-score-details
Resume and Job Description Matching Dataset
Overview
This dataset contains 1,031 samples of resumes and job descriptions (JDs) generated and assessed using GPT-4o. The primary goal of this dataset is to evaluate the alignment between resumes and job descriptions, aiding in the study of resume relevance, skill alignment, and job fit scoring based on predefined criteria.
Dataset Composition
The dataset includes resumes matched with job descriptions, with the… See the full description on the dataset page: https://huggingface.co/datasets/netsol/resume-score-details.ub-networking-dataset-2024-2CDB_DEC2024-CochranSampled_MiniLLMV2_Emb
CrediBench Web Content Embeddings (December 2024)
This repository contains MiniLLM_MultiLingual_V2 embeddings for the CrediBench WebContent (December 2024) dataset.
Source dataset:https://huggingface.co/datasets/Hussein-Abdallah/CrediBench-WebContent-Dec2024_CochranSampled
Overview
The source dataset is a Cochran-sampled subset of the December 2024 CrediBench WebContent corpus. Documents are sampled independently for each web domain using Cochran's sampling… See the full description on the dataset page: https://huggingface.co/datasets/credi-net/CDB_DEC2024-CochranSampled_MiniLLMV2_Emb.instructional_code-search-net-python
Dataset Card for "instructional_code-search-net-python"
Dataset Summary
This is an instructional dataset for Python.
The dataset contains two different kind of tasks:
Given a piece of code generate a description of what it does.
Given a description generate a piece of code that fulfils the description.
Languages
The dataset is in English.
Data Splits
There are no splits.
Dataset Creation
May of 2023
Curation Rationale
This… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-python.network_paper
