datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ids-project-artifactscic-ids-2017
CIC-IDS-2017 Dataset
This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use.
Dataset Structure
Configurations
machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs).
traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC.
Raw Data
The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/bvsam/cic-ids-2017.CIC-IDS2017We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily.
NIDS Datasets
The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/rdpahalavan/CIC-IDS2017.CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles:
Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.the-stack-v2-train-smol-ids-updatedUpdate on The Stack V2 dataset: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids
All repos from original dataset are parsed with Github API and re-downloaded,
so respective updates are kept, metadata is updated. This took 10+ days to process due to GraphQL limits.
Filtering rules
Removed repos with no update in the last 6 years (no updates since September 2019)
Removed files with a single line
Removed repos with a single file
Removed repos with more than 99%… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated.pcod_graph_idsthe-stack-v2-train-smol-ids
The Stack v2
The dataset consists of 4 versions:
bigcode/the-stack-v2: the full "The Stack v2" dataset
bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated
bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories.
bigcode/the-stack-v2-train-smol-ids: based on the bigcode/the-stack-v2-dedup… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids.the-stack-v2-train-full-ids
The Stack v2
The dataset consists of 4 versions:
bigcode/the-stack-v2: the full "The Stack v2" dataset
bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated
bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here
bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.selfies-ids-cleaned32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-32B
with thinking mode off. Each thought is a few dense sentences of reasoning about the next
8 tokens after a cut, written from the document prefix alone — the generator never sees the
continuation. Stored thought_text includes the <thought>/</thought> wrapper.
The 4B parity counterpart is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next
8 tokens after a cut, written from the document prefix alone — the generator never sees the
continuation. Stored thought_text includes the <thought>/</thought> wrapper.
This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.idsse-data
IDSSE — Sportec Open Tracking & Event Data
Mirror of the IDSSE dataset, originally hosted on Figshare, re-hosted here for stable and fast access via kloppy and related tools.
Contents
7 complete matches from the German Bundesliga season 2022/23 (1st and 2nd division), collected by Sportec Solutions using TRACAB optical tracking technology.
Each match includes:
Tracking data — position of all players and the ball at 25 fps
Event data — synchronized match events (passes… See the full description on the dataset page: https://huggingface.co/datasets/pysport/idsse-data.id_scholar
ID Scholar
ID Scholar is a collection of Indonesian scholarly articles or journals compiled from most journals. This dataset is processed by downloading all PDFs and converting them to text using Huridocs' PDF Layout Analysis. This dataset has not been cleaned in any way except for language filtering using FastText.
msa-hotpotqa-qa-with-idsbelka-selfies-idsarxiv-author-affiliations-matched-ror-ids
arXiv Author Affiliations
This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers.
Dataset Description
This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.dml-fl-iot-ids-dynamic3L-K5ids-paragraph-commentary-corpusThe data was collected from GitHub using https://github.com/cheop-byeon/RFCRationaleBuilder.
This is our first version of data for RFC Rationale.
Dataset Citation
If you find this dataset useful and include it in your studies, please cite our paper:
@inproceedings{bian2024tell,
title={Tell Me Why: Language Models Help Explain the Rationale Behind Internet Protocol Design},
author={Bian, Jie and Welzl, Michael and Kutuzov, Andrey and Arefyev, Nikolay},
booktitle={2024 IEEE… See the full description on the dataset page: https://huggingface.co/datasets/jiebi/ids-paragraph-commentary-corpus.CIC-IDS2018LICENSE
You may redistribute, republish, and mirror the CSE-CIC-IDS2018 dataset in any form. However, any use or redistribution of the data must include a citation to the CSE-CIC-IDS2018 dataset and a link to this page in AWS.
Research paper outlining the details of analyzing the similar IDS/IPS dataset and related principles:
Iman Sharafaldin, Arash Habibi Lashkari, and Ali A. Ghorbani, “Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization”, 4th… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2018.dml-fl-iot-ids-static3L4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a few dense sentences that the
generator wrapped in a literal <thought>…</thought> block (plain text, not special tokens).
Only the part after </think> becomes the VALUE, with every <thought> / </thought> tag
removed… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a few dense, declarative sentences (about two to four) inside a
<thought>…</thought> block, focused on the exact state at the cut and what the local
grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.msa-musique-qa-with-idszone-pretrain-ids-datathe-stack-v2-train-smol-ids-updated-contentIncrementally uploaded Parquet shards under data/ with columns:
repo_name: str
text: str
Downloaded all repos from https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated and run
formatting / linting / import sort on all files
Using ruff + black + ty stack for python and biomejs for js / ts / html / css / json / graphql
Total amount of tokens: ~100B
Example:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated-content.selfies-train-idscc-ids-to-title-multilingual
Common Crawl IDs to Titles
Per-config manifests mapping document IDs to their extracted titles.
Usage
from datasets import load_dataset
titles = load_dataset("orionweller/cc-ids-to-title-multilingual-2", "fw-edu")
cic-ids-2017
CIC-IDS-2017 Dataset
This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use.
Dataset Structure
Configurations
machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs).
traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC.
Raw Data
The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/Ariasyah/cic-ids-2017.4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a single paragraph of dense
reasoning about the immediate continuation. Only the part after </think> — the paragraph —
becomes the VALUE. The reasoning inside the think block is dropped.
The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.aflow_graph_ids
