CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01theelderemo /pentesting-explanations Pentesting Explanations - Adversarial Reasoning & Vulnerability Research A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names. The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/pentesting-explanations.texttext-generation10K<n<100K4 likes355 downloads5mo agoHugging Face02penfever /snowball-5.7t-sft-eval-artifacts Snowball 5.7T cold-start SFT evaluation artifacts This dataset archives the evaluation records, sampled traces, resolved launch configurations, analysis inputs, and derived tables for marin-community/marin#8225. The experiment compares the 5.7T-token Snowball cooldown with its Chat, Thinking, and Nemotron-Terminal SFT descendants. It also includes the corresponding 2T-token cooldown cohort. The top-level EVAL_RESULTS.csv in the experiment record is generated from the durable… See the full description on the dataset page: https://huggingface.co/datasets/penfever/snowball-5.7t-sft-eval-artifacts.text-generation0 likes338 downloads1mo agoHugging Face03almanach /penicillin Penicillin dataset Paper: https://arxiv.org/abs/2510.25771 Note This dataset is a processed combination of existing benchmark datasets. The licensing and data rights belong to the original dataset owners. Please refer to the original datasets for more information. tabulartext-generation1M<n<10M0 likes262 downloads9mo agoHugging Face04pe-nlp /ov-kit-doctext-generation0 likes244 downloads2y agoHugging Face05almanach /penicillin_plus Penicillin-Plus dataset Paper: https://arxiv.org/abs/2510.25771 Note This dataset is a processed combination of existing benchmark datasets. The licensing and data rights belong to the original dataset owners. Please refer to the original datasets for more information. tabulartext-generation1M<n<10M1 likes214 downloads9mo agoHugging Face06Lots-of-LoRAs /task149_afs_argument_quality_death_penalty Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task149_afs_argument_quality_death_penalty Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task149_afs_argument_quality_death_penalty.texttext-generation1K<n<10K0 likes153 downloads2y agoHugging Face07me-aas /pentesting-explanations Pentesting Explanations - Adversarial Reasoning & Vulnerability Research A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names. The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/pentesting-explanations.texttext-generation10K<n<100K4 likes144 downloads4mo agoHugging Face08THEATLAS /PENS Dataset Card for PENS: PErsonalized News headlineS PENS is an English dataset for Personalized News Headline Generation Research. It contains two parts for training and test individually. The training set was collected from anonymized user impressions logs of Microsoft News website, and the test set is manually-created by hundreds of native speakers to enable a fair testbed for evaluating models in an offline mode. PENS contains about 113k English news articles whose topics are… See the full description on the dataset page: https://huggingface.co/datasets/THEATLAS/PENS.imagetext-generation100K<n<1M5 likes116 downloads2y agoHugging Face09shaikat005 /medium-web-pentesting Medium Web Pentesting Articles Dataset Description A curated collection of 357 Medium articles focused on web penetration testing, scraped from Medium's search results for the query web pentesting. Each record includes article metadata and the opening snippet of the article body. This dataset is useful for NLP tasks such as topic modeling, text classification, content recommendation, and summarization within the cybersecurity domain. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/shaikat005/medium-web-pentesting.tabulartext-classificationn<1K0 likes109 downloads5mo agoHugging Face10CJJones /Synthetic_PenTest_ReportsThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. 📄 100 Samples of Synthetic Automated Penetration Test Reports This dataset contains 100+ realistic, synthetic penetration testing reportsstructured to simulate professional internal security assessments. Each record models the full flow of a pentest engagement, including: Reconnaissance / Discovery Phase Vulnerability Assessment… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_PenTest_Reports.texttext-classification10K<n<100K3 likes83 downloads7mo agoHugging Face11supersamdev /pentestr1-chat Pentest-R1 Chat — TCO Fine-Tuning Dataset Converted from KHenryAegis/Pentest-R1 into Gemma chat-template JSONL for fine-tuning autonomous penetration testing agents with Unsloth Studio. Dataset summary Field Value Conversations 535 Total messages 28,731 Format JSONL — one {"messages": [...]} per line Chat template Gemma (unsloth/gemma-3n-E4B-it) Token p50 / p95 / p99 / max 2,988 / 6,111 / 7,843 / 8,178 Recommended context_length 8192… See the full description on the dataset page: https://huggingface.co/datasets/supersamdev/pentestr1-chat.texttext-generationn<1K0 likes83 downloads4mo agoHugging Face12Lots-of-LoRAs /task145_afs_argument_similarity_death_penalty Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task145_afs_argument_similarity_death_penalty Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task145_afs_argument_similarity_death_penalty.texttext-generation1K<n<10K0 likes82 downloads2y agoHugging Face13GLauzza /Mille-Pensees-Dataset Mille-Pensées-Dataset Dataset Summary The Mille-Pensées-Dataset is a math reasoning dataset with a 50% french / 50% english composition used to train the Mille-Pensées french reasoning model. The source data comes from the following english math reasoning datasets: s1K-1.1 OpenThoughts3-1.2M OpenR1-Math-220k OpenMathReasoning Nemotron-Post-Training-Dataset-v1 LIMO-v2 DeepMath-103K AM-DeepSeek-R1-0528-Distilled The french reasoning data was obtained by translating the… See the full description on the dataset page: https://huggingface.co/datasets/GLauzza/Mille-Pensees-Dataset.texttext-generation10K<n<100K1 likes79 downloads10mo agoHugging Face14alucent /mirror-pentesting-explanationsgated Pentesting Explanations - Adversarial Reasoning & Vulnerability Research A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names. The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-pentesting-explanations.texttext-generation10K<n<100K1 likes77 downloads2mo agoHugging Face15Lots-of-LoRAs /task1167_penn_treebank_coarse_pos_tagging Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1167_penn_treebank_coarse_pos_tagging Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1167_penn_treebank_coarse_pos_tagging.texttext-generation1K<n<10K1 likes71 downloads2y agoHugging Face16louisbrulenaudet /code-pensions-civiles-militaires-retraite Code des pensions civiles et militaires de retraite, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-pensions-civiles-militaires-retraite.tabulartext-generationn<1K0 likes69 downloads2y agoHugging Face17AshishFugare /Pentesting_Datasettexttext-generationn<1K1 likes45 downloads8mo agoHugging Face18mikoube /pentest Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mikoube/pentest.text-generation1K<n<10K0 likes39 downloads3y agoHugging Face19louisbrulenaudet /code-penal Code pénal, non-instruct (2025-05-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models based… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-penal.tabulartext-generation1K<n<10K1 likes38 downloads1y agoHugging Face20pensieves /PreferenceTravelPlanner PreferenceTravelPlanner Dataset PreferenceTravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints and preferences. For more details, see our paper. It is created by augmenting TravelPlanner (See paper for more details) with several common type of preferences under various preference paradigms. Introduction In PreferenceTravelPlanner, for a given query, language agents are expected to formulate a… See the full description on the dataset page: https://huggingface.co/datasets/pensieves/PreferenceTravelPlanner.tabulartext-generation1K<n<10K0 likes34 downloads7mo agoHugging Face21louisbrulenaudet /code-procedure-penale Code de procédure pénale, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-procedure-penale.tabulartext-generation1K<n<10K1 likes31 downloads2y agoHugging Face22louisbrulenaudet /code-justice-penale-mineurs Code de la justice pénale des mineurs, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-justice-penale-mineurs.tabulartext-generationn<1K0 likes30 downloads1y agoHugging Face23pencil-code /anonymized_datagated PencilCode Program Traces Paper: Modeling Student Learning with 3.8 Million Program TracesCode: meghabyte/pencilcode-publicContact: megha@cs.stanford.edu, alexisro@mit.edu, jjb@eng.ufl.edu, jda@mit.edu Gated Dataset — Manual Approval Required.Access is granted pending review by the Pencil Code team. Please submit your request above with a brief description of your intended research use. Any requests will be reviewed by Jeremiah Blanchard at jjb@eng.ufl.edu Dataset… See the full description on the dataset page: https://huggingface.co/datasets/pencil-code/anonymized_data.texttext-generation1M<n<10M3 likes30 downloads4mo agoHugging Face24pendingremove32894 /cantonesewiki_doyouknow Cantonese Question Dataset from Yue Wiki A collection of questions in Cantonese, extracted from Yue Wiki. This dataset contains a variety of questions covering different topics and domains. Disclaimer The content and opinions expressed in this dataset do not represent the views, beliefs, or positions of the dataset creators, contributors, or hosting organizations. This dataset is provided solely for the purpose of improving AI systems' understanding of the Cantonese… See the full description on the dataset page: https://huggingface.co/datasets/pendingremove32894/cantonesewiki_doyouknow.texttext-generation1K<n<10K3 likes29 downloads2y agoHugging Face25Nanthasit /food-penguin-v1 Food Penguin v1 Food Penguin v1 is a compact, domain-specific text-generation dataset from the SakThai family. It provides conversation turns paired with tool definitions and expected function calls in JSON format, optimized for instruction-following, food-domain QA, and lightweight tool-use model training. Dataset Summary Food Penguin v1 is designed for: Instruction-following models Food & restaurant domain QA Tool-use and function-calling training… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/food-penguin-v1.text-generationn<1K0 likes26 downloads2mo agoHugging Face26iradukunda-dev /offences_and_penalties_in_general_2018_datasettexttext-classificationn<1K0 likes24 downloads9mo agoHugging Face27datamarketinglabs /pentadrive-v1 TrueHuman PentaDrive Dataset summary TrueHuman PentaDrive is a compact behavioral matrix for affective and conversational modeling. It organizes human-relevant motivational structure into five drives (coded S, K, A, M, G), each implemented as a set of nodes (stable behavioral motifs). Every node defines three phases—anticipation, release, and block—with: Markers: short textual cues for lightweight classification or retrieval Response kernels: structured hints for… See the full description on the dataset page: https://huggingface.co/datasets/datamarketinglabs/pentadrive-v1.texttext-classification1K<n<10K0 likes24 downloads6mo agoHugging Face28penfever /MM-MathInstruct-longest-20k-solutions-with-images MM-MathInstruct Longest 20K Solutions with Images This dataset contains the top 20,000 samples from MathLLMs/MM-MathInstruct selected by solution length, filtered to include only samples with valid images. Dataset Structure Usage Source This dataset is derived from MathLLMs/MM-MathInstruct by selecting the 20,000 samples with the longest solution text. License Apache 2.0 (inherited from source dataset) textvisual-question-answering10K<n<100K0 likes21 downloads1y agoHugging Face29louisbrulenaudet /code-disciplinaire-penal-marine-marchande Code disciplinaire et pénal de la marine marchande, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-disciplinaire-penal-marine-marchande.tabulartext-generationn<1K0 likes20 downloads1y agoHugging Face30louisbrulenaudet /code-penitentiaire Code pénitentiaire, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-penitentiaire.tabulartext-generation1K<n<10K0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.