CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01theelderemo /pentesting-explanations Pentesting Explanations - Adversarial Reasoning & Vulnerability Research A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names. The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/pentesting-explanations.texttext-generation10K<n<100K4 likes343 downloads5mo agoHugging Face02penfever /snowball-5.7t-sft-eval-artifacts Snowball 5.7T cold-start SFT evaluation artifacts This dataset archives the evaluation records, sampled traces, resolved launch configurations, analysis inputs, and derived tables for marin-community/marin#8225. The experiment compares the 5.7T-token Snowball cooldown with its Chat, Thinking, and Nemotron-Terminal SFT descendants. It also includes the corresponding 2T-token cooldown cohort. The top-level EVAL_RESULTS.csv in the experiment record is generated from the durable… See the full description on the dataset page: https://huggingface.co/datasets/penfever/snowball-5.7t-sft-eval-artifacts.text-generation0 likes339 downloads1mo agoHugging Face03almanach /penicillin Penicillin dataset Paper: https://arxiv.org/abs/2510.25771 Note This dataset is a processed combination of existing benchmark datasets. The licensing and data rights belong to the original dataset owners. Please refer to the original datasets for more information. tabulartext-generation1M<n<10M0 likes269 downloads9mo agoHugging Face04pe-nlp /ov-kit-doctext-generation0 likes244 downloads2y agoHugging Face05almanach /penicillin_plus Penicillin-Plus dataset Paper: https://arxiv.org/abs/2510.25771 Note This dataset is a processed combination of existing benchmark datasets. The licensing and data rights belong to the original dataset owners. Please refer to the original datasets for more information. tabulartext-generation1M<n<10M1 likes216 downloads9mo agoHugging Face06Lots-of-LoRAs /task149_afs_argument_quality_death_penalty Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task149_afs_argument_quality_death_penalty Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task149_afs_argument_quality_death_penalty.texttext-generation1K<n<10K0 likes148 downloads2y agoHugging Face07me-aas /pentesting-explanations Pentesting Explanations - Adversarial Reasoning & Vulnerability Research A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names. The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/pentesting-explanations.texttext-generation10K<n<100K4 likes145 downloads4mo agoHugging Face08THEATLAS /PENS Dataset Card for PENS: PErsonalized News headlineS PENS is an English dataset for Personalized News Headline Generation Research. It contains two parts for training and test individually. The training set was collected from anonymized user impressions logs of Microsoft News website, and the test set is manually-created by hundreds of native speakers to enable a fair testbed for evaluating models in an offline mode. PENS contains about 113k English news articles whose topics are… See the full description on the dataset page: https://huggingface.co/datasets/THEATLAS/PENS.imagetext-generation100K<n<1M5 likes117 downloads2y agoHugging Face09shaikat005 /medium-web-pentesting Medium Web Pentesting Articles Dataset Description A curated collection of 357 Medium articles focused on web penetration testing, scraped from Medium's search results for the query web pentesting. Each record includes article metadata and the opening snippet of the article body. This dataset is useful for NLP tasks such as topic modeling, text classification, content recommendation, and summarization within the cybersecurity domain. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/shaikat005/medium-web-pentesting.tabulartext-classificationn<1K0 likes111 downloads5mo agoHugging Face10supersamdev /pentestr1-chat Pentest-R1 Chat — TCO Fine-Tuning Dataset Converted from KHenryAegis/Pentest-R1 into Gemma chat-template JSONL for fine-tuning autonomous penetration testing agents with Unsloth Studio. Dataset summary Field Value Conversations 535 Total messages 28,731 Format JSONL — one {"messages": [...]} per line Chat template Gemma (unsloth/gemma-3n-E4B-it) Token p50 / p95 / p99 / max 2,988 / 6,111 / 7,843 / 8,178 Recommended context_length 8192… See the full description on the dataset page: https://huggingface.co/datasets/supersamdev/pentestr1-chat.texttext-generationn<1K0 likes81 downloads4mo agoHugging Face11CJJones /Synthetic_PenTest_ReportsThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. 📄 100 Samples of Synthetic Automated Penetration Test Reports This dataset contains 100+ realistic, synthetic penetration testing reportsstructured to simulate professional internal security assessments. Each record models the full flow of a pentest engagement, including: Reconnaissance / Discovery Phase Vulnerability Assessment… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_PenTest_Reports.texttext-classification10K<n<100K3 likes80 downloads7mo agoHugging Face12GLauzza /Mille-Pensees-Dataset Mille-Pensées-Dataset Dataset Summary The Mille-Pensées-Dataset is a math reasoning dataset with a 50% french / 50% english composition used to train the Mille-Pensées french reasoning model. The source data comes from the following english math reasoning datasets: s1K-1.1 OpenThoughts3-1.2M OpenR1-Math-220k OpenMathReasoning Nemotron-Post-Training-Dataset-v1 LIMO-v2 DeepMath-103K AM-DeepSeek-R1-0528-Distilled The french reasoning data was obtained by translating the… See the full description on the dataset page: https://huggingface.co/datasets/GLauzza/Mille-Pensees-Dataset.texttext-generation10K<n<100K1 likes80 downloads10mo agoHugging Face13alucent /mirror-pentesting-explanationsgated Pentesting Explanations - Adversarial Reasoning & Vulnerability Research A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names. The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-pentesting-explanations.texttext-generation10K<n<100K1 likes79 downloads2mo agoHugging Face14Lots-of-LoRAs /task145_afs_argument_similarity_death_penalty Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task145_afs_argument_similarity_death_penalty Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task145_afs_argument_similarity_death_penalty.texttext-generation1K<n<10K0 likes77 downloads2y agoHugging Face15louisbrulenaudet /code-pensions-civiles-militaires-retraite Code des pensions civiles et militaires de retraite, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-pensions-civiles-militaires-retraite.tabulartext-generationn<1K0 likes69 downloads2y agoHugging Face16Lots-of-LoRAs /task1167_penn_treebank_coarse_pos_tagging Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1167_penn_treebank_coarse_pos_tagging Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1167_penn_treebank_coarse_pos_tagging.texttext-generation1K<n<10K1 likes66 downloads2y agoHugging Face17AshishFugare /Pentesting_Datasettexttext-generationn<1K1 likes44 downloads8mo agoHugging Face18mikoube /pentest Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mikoube/pentest.text-generation1K<n<10K0 likes38 downloads3y agoHugging Face19louisbrulenaudet /code-penal Code pénal, non-instruct (2025-05-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models based… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-penal.tabulartext-generation1K<n<10K1 likes36 downloads1y agoHugging Face20louisbrulenaudet /code-procedure-penale Code de procédure pénale, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-procedure-penale.tabulartext-generation1K<n<10K1 likes32 downloads2y agoHugging Face21louisbrulenaudet /code-justice-penale-mineurs Code de la justice pénale des mineurs, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-justice-penale-mineurs.tabulartext-generationn<1K0 likes31 downloads1y agoHugging Face22pendingremove32894 /cantonesewiki_doyouknow Cantonese Question Dataset from Yue Wiki A collection of questions in Cantonese, extracted from Yue Wiki. This dataset contains a variety of questions covering different topics and domains. Disclaimer The content and opinions expressed in this dataset do not represent the views, beliefs, or positions of the dataset creators, contributors, or hosting organizations. This dataset is provided solely for the purpose of improving AI systems' understanding of the Cantonese… See the full description on the dataset page: https://huggingface.co/datasets/pendingremove32894/cantonesewiki_doyouknow.texttext-generation1K<n<10K3 likes30 downloads2y agoHugging Face23pencil-code /anonymized_datagated PencilCode Program Traces Paper: Modeling Student Learning with 3.8 Million Program TracesCode: meghabyte/pencilcode-publicContact: megha@cs.stanford.edu, alexisro@mit.edu, jjb@eng.ufl.edu, jda@mit.edu Gated Dataset — Manual Approval Required.Access is granted pending review by the Pencil Code team. Please submit your request above with a brief description of your intended research use. Any requests will be reviewed by Jeremiah Blanchard at jjb@eng.ufl.edu Dataset… See the full description on the dataset page: https://huggingface.co/datasets/pencil-code/anonymized_data.texttext-generation1M<n<10M3 likes27 downloads4mo agoHugging Face24iradukunda-dev /offences_and_penalties_in_general_2018_datasettexttext-classificationn<1K0 likes26 downloads9mo agoHugging Face25datamarketinglabs /pentadrive-v1 TrueHuman PentaDrive Dataset summary TrueHuman PentaDrive is a compact behavioral matrix for affective and conversational modeling. It organizes human-relevant motivational structure into five drives (coded S, K, A, M, G), each implemented as a set of nodes (stable behavioral motifs). Every node defines three phases—anticipation, release, and block—with: Markers: short textual cues for lightweight classification or retrieval Response kernels: structured hints for… See the full description on the dataset page: https://huggingface.co/datasets/datamarketinglabs/pentadrive-v1.texttext-classification1K<n<10K0 likes26 downloads6mo agoHugging Face26pensieves /PreferenceTravelPlanner PreferenceTravelPlanner Dataset PreferenceTravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints and preferences. For more details, see our paper. It is created by augmenting TravelPlanner (See paper for more details) with several common type of preferences under various preference paradigms. Introduction In PreferenceTravelPlanner, for a given query, language agents are expected to formulate a… See the full description on the dataset page: https://huggingface.co/datasets/pensieves/PreferenceTravelPlanner.tabulartext-generation1K<n<10K0 likes25 downloads7mo agoHugging Face27penfever /MM-MathInstruct-longest-20k-solutions-with-images MM-MathInstruct Longest 20K Solutions with Images This dataset contains the top 20,000 samples from MathLLMs/MM-MathInstruct selected by solution length, filtered to include only samples with valid images. Dataset Structure Usage Source This dataset is derived from MathLLMs/MM-MathInstruct by selecting the 20,000 samples with the longest solution text. License Apache 2.0 (inherited from source dataset) textvisual-question-answering10K<n<100K0 likes22 downloads1y agoHugging Face28louisbrulenaudet /code-disciplinaire-penal-marine-marchande Code disciplinaire et pénal de la marine marchande, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-disciplinaire-penal-marine-marchande.tabulartext-generationn<1K0 likes20 downloads1y agoHugging Face29louisbrulenaudet /code-penitentiaire Code pénitentiaire, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-penitentiaire.tabulartext-generation1K<n<10K0 likes20 downloads2y agoHugging Face30Nanthasit /food-penguin-v1 Food Penguin v1 Food Penguin v1 is a compact, domain-specific text-generation dataset from the SakThai family. It provides conversation turns paired with tool definitions and expected function calls in JSON format, optimized for instruction-following, food-domain QA, and lightweight tool-use model training. Dataset Summary Food Penguin v1 is designed for: Instruction-following models Food & restaurant domain QA Tool-use and function-calling training… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/food-penguin-v1.text-generationn<1K0 likes19 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.