CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bvsam /cic-ids-2017 CIC-IDS-2017 Dataset This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use. Dataset Structure Configurations machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs). traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC. Raw Data The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/bvsam/cic-ids-2017.tabulartabular-classification1M<n<10M3 likes1.9k downloads10mo agoHugging Face02c01dsnap /CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.tabular1M<n<10M5 likes1.2k downloads3y agoHugging Face03thepowerfuldeez /the-stack-v2-train-smol-ids-updatedUpdate on The Stack V2 dataset: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids All repos from original dataset are parsed with Github API and re-downloaded, so respective updates are kept, metadata is updated. This took 10+ days to process due to GraphQL limits. Filtering rules Removed repos with no update in the last 6 years (no updates since September 2019) Removed files with a single line Removed repos with a single file Removed repos with more than 99%… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated.tabular100K<n<1M0 likes951 downloads1y agoHugging Face04bigcode /the-stack-v2-train-smol-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. bigcode/the-stack-v2-train-smol-ids: based on the bigcode/the-stack-v2-dedup… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids.tabulartext-generation10M<n<100M60 likes823 downloads2mo agoHugging Face05bigcode /the-stack-v2-train-full-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.tabulartext-generation10M<n<100M65 likes665 downloads2mo agoHugging Face06HoangHa /selfies-ids-cleanedtext100M<n<1B0 likes615 downloads2y agoHugging Face07JackHsieh /32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-32B with thinking mode off. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. The 4B parity counterpart is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes591 downloads23d agoHugging Face08JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes518 downloads7d agoHugging Face09Wikidepia /id_scholar ID Scholar ID Scholar is a collection of Indonesian scholarly articles or journals compiled from most journals. This dataset is processed by downloading all PDFs and converting them to text using Huridocs' PDF Layout Analysis. This dataset has not been cleaned in any way except for language filtering using FastText. text100K<n<1M1 likes470 downloads2y agoHugging Face10cometadata /arxiv-author-affiliations-matched-ror-ids arXiv Author Affiliations This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers. Dataset Description This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.texttext-classification1M<n<10M1 likes426 downloads8mo agoHugging Face11HoangHa /belka-selfies-idstext10M<n<100M0 likes423 downloads2y agoHugging Face12ragrawal36 /msa-hotpotqa-qa-with-idstext1K<n<10K0 likes423 downloads5mo agoHugging Face13thepowerfuldeez /the-stack-v2-train-smol-ids-updated-contentIncrementally uploaded Parquet shards under data/ with columns: repo_name: str text: str Downloaded all repos from https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated and run formatting / linting / import sort on all files Using ruff + black + ty stack for python and biomejs for js / ts / html / css / json / graphql Total amount of tokens: ~100B Example: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated-content.text100K<n<1M0 likes356 downloads1y agoHugging Face14JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a few dense sentences that the generator wrapped in a literal <thought>…</thought> block (plain text, not special tokens). Only the part after </think> becomes the VALUE, with every <thought> / </thought> tag removed… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes331 downloads16d agoHugging Face15JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a few dense, declarative sentences (about two to four) inside a <thought>…</thought> block, focused on the exact state at the cut and what the local grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes329 downloads16d agoHugging Face16HoangHa /selfies-train-idstext100M<n<1B0 likes320 downloads2y agoHugging Face17orionweller /cc-ids-to-title-multilingual Common Crawl IDs to Titles Per-config manifests mapping document IDs to their extracted titles. Usage from datasets import load_dataset titles = load_dataset("orionweller/cc-ids-to-title-multilingual-2", "fw-edu") text100M<n<1B0 likes310 downloads6mo agoHugging Face18Ariasyah /cic-ids-2017 CIC-IDS-2017 Dataset This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use. Dataset Structure Configurations machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs). traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC. Raw Data The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/Ariasyah/cic-ids-2017.tabulartabular-classification1M<n<10M0 likes303 downloads9mo agoHugging Face19JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a single paragraph of dense reasoning about the immediate continuation. Only the part after </think> — the paragraph — becomes the VALUE. The reasoning inside the think block is dropped. The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes302 downloads16d agoHugging Face20ragrawal36 /msa-musique-qa-with-idstextn<1K0 likes301 downloads5mo agoHugging Face21JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes286 downloads16d agoHugging Face22ragrawal36 /msa-hotpotqa-docs-with-idstext1K<n<10K0 likes276 downloads5mo agoHugging Face23veera33 /CIC-IDS2017We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily. NIDS Datasets The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/veera33/CIC-IDS2017.tabulartext-classification100M<n<1B0 likes270 downloads7mo agoHugging Face24Wikidepia /id_scholar_huridocstext1M<n<10M0 likes269 downloads2y agoHugging Face25vibhuiitj /UltraData-Math-L2-preview-embedded-with-idstabular10M<n<100M0 likes263 downloads6mo agoHugging Face26JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes253 downloads8d agoHugging Face27underfrog /msa-2wikimultihopqa-qa-with-idstext1K<n<10K0 likes251 downloads3mo agoHugging Face28lihaoxin2020 /abstractive-content-based-IDs Abstractive Content-Based Document IDs for Generative Retrieval Dataset for Summarization-Based Document IDs for Generative Retrieval with Language Models. Update [03/04/2025] Upload validation and test set of ACID. Add tokenized subset. @misc{li2024summarizationbaseddocumentidsgenerative, title={Summarization-Based Document IDs for Generative Retrieval with Language Models}, author={Haoxin Li and Daniel Cheng and Phillip Keung and Jungo Kasai and Noah A.… See the full description on the dataset page: https://huggingface.co/datasets/lihaoxin2020/abstractive-content-based-IDs.text1M<n<10M0 likes224 downloads2y agoHugging Face29pcy12345BSU /CIC-IDS-2017We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily. NIDS Datasets The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/pcy12345BSU/CIC-IDS-2017.tabulartext-classification100M<n<1B1 likes203 downloads6mo agoHugging Face30JackHsieh /luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled Every non-first chunk of every document carries a thought: the gpt-5.6-luna reasoning thought where one was generated, and a content-free pause thought everywhere else. luna chunk <|reserved_special_token_1|> luna reasoning <|reserved_special_token_2|> filler chunk <|reserved_special_token_1|> 256x <|reserved_special_token_0|> <|reserved_special_token_2|> The filler is 258 tokens. Chunk 0 is excluded… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled.tabulartext-generation1M<n<10M0 likes189 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.