unbalance
DOS201X_highly_unbalanced
DoS201X Highly Unbalanced (Federated, Pre-Partitioned)
This dataset has been specifically prepared for federated learning (FL) experiments on DDoS attack detection.
It is partitioned among 15 clients, with each client assigned a single attack type together with a portion of benign traffic,
resulting in a non-IID distribution of network flows across clients. This partitioning is designed to reproduce a realistic
FL setting in which individual clients have access to different… See the full description on the dataset page: https://huggingface.co/datasets/silviocretti/DOS201X_highly_unbalanced.ImageNet15_animals_unbalanced_aug1
Dataset Card for "ImageNet15_animals_unbalanced_aug1"
More Information needed
DOS2019_highly_unbalanced
CIC-DDoS2019 Highly Unbalanced (Federated, Pre-Partitioned)
This dataset contains a collection of DDoS attacks. It is a preprocessed, repartitioned derivative of the CIC-DDoS2019
dataset, originally published by the Canadian Institute for Cybersecurity (CIC),
University of New Brunswick.
More details about the CIC-DDoS2019 dataset can be found on this page
and in the following scientific paper:
Iman Sharafaldin, Arash Habibi Lashkari, Saqib Hakak, and Ali A. Ghorbani… See the full description on the dataset page: https://huggingface.co/datasets/silviocretti/DOS2019_highly_unbalanced.unbalanced_d4rlAudioSet_unbalanced_video200k-tulu-2-unbalanced
Tulu 2 Unfiltered - 200k subset
This the 200k subset of the 'unfiltered' version of the Tulu v2 SFT mixture, created by collating the original Tulu 2 sources and avoiding downsampling.
This was used for the 200k-size experiments.
Details
The dataset consists of a mix of :
FLAN (Apache 2.0, we only sample 961,322 samples along with 398,439 CoT samples from the full set for this data pool)
Open Assistant 1 (Apache 2.0)
ShareGPT (Apache 2.0 listed, no official repo… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/200k-tulu-2-unbalanced.
