CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hossam3759180 /ids-project-artifactsimage10M<n<100M0 likes4k downloads2d agoHugging Face02bvsam /cic-ids-2017 CIC-IDS-2017 Dataset This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use. Dataset Structure Configurations machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs). traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC. Raw Data The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/bvsam/cic-ids-2017.tabulartabular-classification1M<n<10M3 likes1.7k downloads10mo agoHugging Face03rdpahalavan /CIC-IDS2017We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily. NIDS Datasets The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/rdpahalavan/CIC-IDS2017.text-classification100M<n<1B4 likes1.5k downloads3y agoHugging Face04c01dsnap /CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.tabular1M<n<10M5 likes1.2k downloads3y agoHugging Face05thepowerfuldeez /the-stack-v2-train-smol-ids-updatedUpdate on The Stack V2 dataset: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids All repos from original dataset are parsed with Github API and re-downloaded, so respective updates are kept, metadata is updated. This took 10+ days to process due to GraphQL limits. Filtering rules Removed repos with no update in the last 6 years (no updates since September 2019) Removed files with a single line Removed repos with a single file Removed repos with more than 99%… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated.tabular100K<n<1M0 likes932 downloads1y agoHugging Face06kamabata /pcod_graph_ids0 likes839 downloads1y agoHugging Face07bigcode /the-stack-v2-train-smol-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. bigcode/the-stack-v2-train-smol-ids: based on the bigcode/the-stack-v2-dedup… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids.tabulartext-generation10M<n<100M60 likes780 downloads2mo agoHugging Face08bigcode /the-stack-v2-train-full-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.tabulartext-generation10M<n<100M65 likes612 downloads2mo agoHugging Face09HoangHa /selfies-ids-cleanedtext100M<n<1B0 likes606 downloads2y agoHugging Face10JackHsieh /32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-32B with thinking mode off. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. The 4B parity counterpart is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes584 downloads22d agoHugging Face11JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes540 downloads6d agoHugging Face12pysport /idsse-data IDSSE — Sportec Open Tracking & Event Data Mirror of the IDSSE dataset, originally hosted on Figshare, re-hosted here for stable and fast access via kloppy and related tools. Contents 7 complete matches from the German Bundesliga season 2022/23 (1st and 2nd division), collected by Sportec Solutions using TRACAB optical tracking technology. Each match includes: Tracking data — position of all players and the ball at 25 fps Event data — synchronized match events (passes… See the full description on the dataset page: https://huggingface.co/datasets/pysport/idsse-data.geospatialother1 likes500 downloads7mo agoHugging Face13Wikidepia /id_scholar ID Scholar ID Scholar is a collection of Indonesian scholarly articles or journals compiled from most journals. This dataset is processed by downloading all PDFs and converting them to text using Huridocs' PDF Layout Analysis. This dataset has not been cleaned in any way except for language filtering using FastText. text100K<n<1M1 likes479 downloads2y agoHugging Face14ragrawal36 /msa-hotpotqa-qa-with-idstext1K<n<10K0 likes440 downloads5mo agoHugging Face15HoangHa /belka-selfies-idstext10M<n<100M0 likes425 downloads2y agoHugging Face16cometadata /arxiv-author-affiliations-matched-ror-ids arXiv Author Affiliations This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers. Dataset Description This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.texttext-classification1M<n<10M1 likes416 downloads8mo agoHugging Face17JabaleNurAdnan /dml-fl-iot-ids-dynamic3L-K50 likes365 downloads4mo agoHugging Face18jiebi /ids-paragraph-commentary-corpusThe data was collected from GitHub using https://github.com/cheop-byeon/RFCRationaleBuilder. This is our first version of data for RFC Rationale. Dataset Citation If you find this dataset useful and include it in your studies, please cite our paper: @inproceedings{bian2024tell, title={Tell Me Why: Language Models Help Explain the Rationale Behind Internet Protocol Design}, author={Bian, Jie and Welzl, Michael and Kutuzov, Andrey and Arefyev, Nikolay}, booktitle={2024 IEEE… See the full description on the dataset page: https://huggingface.co/datasets/jiebi/ids-paragraph-commentary-corpus.text-classification0 likes351 downloads5mo agoHugging Face19c01dsnap /CIC-IDS2018LICENSE You may redistribute, republish, and mirror the CSE-CIC-IDS2018 dataset in any form. However, any use or redistribution of the data must include a citation to the CSE-CIC-IDS2018 dataset and a link to this page in AWS. Research paper outlining the details of analyzing the similar IDS/IPS dataset and related principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A. Ghorbani, “Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization”, 4th… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2018.0 likes349 downloads3y agoHugging Face20JabaleNurAdnan /dml-fl-iot-ids-static3L0 likes338 downloads4mo agoHugging Face21JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a few dense sentences that the generator wrapped in a literal <thought>…</thought> block (plain text, not special tokens). Only the part after </think> becomes the VALUE, with every <thought> / </thought> tag removed… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes330 downloads15d agoHugging Face22JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a few dense, declarative sentences (about two to four) inside a <thought>…</thought> block, focused on the exact state at the cut and what the local grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes329 downloads15d agoHugging Face23ragrawal36 /msa-musique-qa-with-idstextn<1K0 likes326 downloads5mo agoHugging Face24sullivan1502 /zone-pretrain-ids-data1M<n<10M0 likes315 downloads2mo agoHugging Face25thepowerfuldeez /the-stack-v2-train-smol-ids-updated-contentIncrementally uploaded Parquet shards under data/ with columns: repo_name: str text: str Downloaded all repos from https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated and run formatting / linting / import sort on all files Using ruff + black + ty stack for python and biomejs for js / ts / html / css / json / graphql Total amount of tokens: ~100B Example: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated-content.text100K<n<1M0 likes314 downloads1y agoHugging Face26HoangHa /selfies-train-idstext100M<n<1B0 likes313 downloads2y agoHugging Face27orionweller /cc-ids-to-title-multilingual Common Crawl IDs to Titles Per-config manifests mapping document IDs to their extracted titles. Usage from datasets import load_dataset titles = load_dataset("orionweller/cc-ids-to-title-multilingual-2", "fw-edu") text100M<n<1B0 likes307 downloads6mo agoHugging Face28Ariasyah /cic-ids-2017 CIC-IDS-2017 Dataset This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use. Dataset Structure Configurations machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs). traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC. Raw Data The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/Ariasyah/cic-ids-2017.tabulartabular-classification1M<n<10M0 likes302 downloads9mo agoHugging Face29JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a single paragraph of dense reasoning about the immediate continuation. Only the part after </think> — the paragraph — becomes the VALUE. The reasoning inside the think block is dropped. The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes302 downloads15d agoHugging Face30kamabata /aflow_graph_ids0 likes289 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.