CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01theelderemo /pentesting-explanations Pentesting Explanations - Adversarial Reasoning & Vulnerability Research A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names. The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/pentesting-explanations.texttext-generation10K<n<100K4 likes343 downloads5mo agoHugging Face02almanach /penicillin Penicillin dataset Paper: https://arxiv.org/abs/2510.25771 Note This dataset is a processed combination of existing benchmark datasets. The licensing and data rights belong to the original dataset owners. Please refer to the original datasets for more information. tabulartext-generation1M<n<10M0 likes269 downloads9mo agoHugging Face03almanach /penicillin_plus Penicillin-Plus dataset Paper: https://arxiv.org/abs/2510.25771 Note This dataset is a processed combination of existing benchmark datasets. The licensing and data rights belong to the original dataset owners. Please refer to the original datasets for more information. tabulartext-generation1M<n<10M1 likes216 downloads9mo agoHugging Face04Lots-of-LoRAs /task149_afs_argument_quality_death_penalty Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task149_afs_argument_quality_death_penalty Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task149_afs_argument_quality_death_penalty.texttext-generation1K<n<10K0 likes148 downloads2y agoHugging Face05me-aas /pentesting-explanations Pentesting Explanations - Adversarial Reasoning & Vulnerability Research A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names. The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/pentesting-explanations.texttext-generation10K<n<100K4 likes145 downloads4mo agoHugging Face06shaikat005 /medium-web-pentesting Medium Web Pentesting Articles Dataset Description A curated collection of 357 Medium articles focused on web penetration testing, scraped from Medium's search results for the query web pentesting. Each record includes article metadata and the opening snippet of the article body. This dataset is useful for NLP tasks such as topic modeling, text classification, content recommendation, and summarization within the cybersecurity domain. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/shaikat005/medium-web-pentesting.tabulartext-classificationn<1K0 likes111 downloads5mo agoHugging Face07supersamdev /pentestr1-chat Pentest-R1 Chat — TCO Fine-Tuning Dataset Converted from KHenryAegis/Pentest-R1 into Gemma chat-template JSONL for fine-tuning autonomous penetration testing agents with Unsloth Studio. Dataset summary Field Value Conversations 535 Total messages 28,731 Format JSONL — one {"messages": [...]} per line Chat template Gemma (unsloth/gemma-3n-E4B-it) Token p50 / p95 / p99 / max 2,988 / 6,111 / 7,843 / 8,178 Recommended context_length 8192… See the full description on the dataset page: https://huggingface.co/datasets/supersamdev/pentestr1-chat.texttext-generationn<1K0 likes81 downloads4mo agoHugging Face08CJJones /Synthetic_PenTest_ReportsThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. 📄 100 Samples of Synthetic Automated Penetration Test Reports This dataset contains 100+ realistic, synthetic penetration testing reportsstructured to simulate professional internal security assessments. Each record models the full flow of a pentest engagement, including: Reconnaissance / Discovery Phase Vulnerability Assessment… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_PenTest_Reports.texttext-classification10K<n<100K3 likes80 downloads7mo agoHugging Face09GLauzza /Mille-Pensees-Dataset Mille-Pensées-Dataset Dataset Summary The Mille-Pensées-Dataset is a math reasoning dataset with a 50% french / 50% english composition used to train the Mille-Pensées french reasoning model. The source data comes from the following english math reasoning datasets: s1K-1.1 OpenThoughts3-1.2M OpenR1-Math-220k OpenMathReasoning Nemotron-Post-Training-Dataset-v1 LIMO-v2 DeepMath-103K AM-DeepSeek-R1-0528-Distilled The french reasoning data was obtained by translating the… See the full description on the dataset page: https://huggingface.co/datasets/GLauzza/Mille-Pensees-Dataset.texttext-generation10K<n<100K1 likes80 downloads10mo agoHugging Face10alucent /mirror-pentesting-explanationsgated Pentesting Explanations - Adversarial Reasoning & Vulnerability Research A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names. The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-pentesting-explanations.texttext-generation10K<n<100K1 likes79 downloads2mo agoHugging Face11Lots-of-LoRAs /task145_afs_argument_similarity_death_penalty Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task145_afs_argument_similarity_death_penalty Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task145_afs_argument_similarity_death_penalty.texttext-generation1K<n<10K0 likes77 downloads2y agoHugging Face12louisbrulenaudet /code-pensions-civiles-militaires-retraite Code des pensions civiles et militaires de retraite, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-pensions-civiles-militaires-retraite.tabulartext-generationn<1K0 likes69 downloads2y agoHugging Face13Lots-of-LoRAs /task1167_penn_treebank_coarse_pos_tagging Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1167_penn_treebank_coarse_pos_tagging Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1167_penn_treebank_coarse_pos_tagging.texttext-generation1K<n<10K1 likes66 downloads2y agoHugging Face14AshishFugare /Pentesting_Datasettexttext-generationn<1K1 likes44 downloads8mo agoHugging Face15louisbrulenaudet /code-penal Code pénal, non-instruct (2025-05-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models based… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-penal.tabulartext-generation1K<n<10K1 likes36 downloads1y agoHugging Face16louisbrulenaudet /code-procedure-penale Code de procédure pénale, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-procedure-penale.tabulartext-generation1K<n<10K1 likes32 downloads2y agoHugging Face17louisbrulenaudet /code-justice-penale-mineurs Code de la justice pénale des mineurs, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-justice-penale-mineurs.tabulartext-generationn<1K0 likes31 downloads1y agoHugging Face18pendingremove32894 /cantonesewiki_doyouknow Cantonese Question Dataset from Yue Wiki A collection of questions in Cantonese, extracted from Yue Wiki. This dataset contains a variety of questions covering different topics and domains. Disclaimer The content and opinions expressed in this dataset do not represent the views, beliefs, or positions of the dataset creators, contributors, or hosting organizations. This dataset is provided solely for the purpose of improving AI systems' understanding of the Cantonese… See the full description on the dataset page: https://huggingface.co/datasets/pendingremove32894/cantonesewiki_doyouknow.texttext-generation1K<n<10K3 likes30 downloads2y agoHugging Face19pencil-code /anonymized_datagated PencilCode Program Traces Paper: Modeling Student Learning with 3.8 Million Program TracesCode: meghabyte/pencilcode-publicContact: megha@cs.stanford.edu, alexisro@mit.edu, jjb@eng.ufl.edu, jda@mit.edu Gated Dataset — Manual Approval Required.Access is granted pending review by the Pencil Code team. Please submit your request above with a brief description of your intended research use. Any requests will be reviewed by Jeremiah Blanchard at jjb@eng.ufl.edu Dataset… See the full description on the dataset page: https://huggingface.co/datasets/pencil-code/anonymized_data.texttext-generation1M<n<10M3 likes27 downloads4mo agoHugging Face20iradukunda-dev /offences_and_penalties_in_general_2018_datasettexttext-classificationn<1K0 likes26 downloads9mo agoHugging Face21datamarketinglabs /pentadrive-v1 TrueHuman PentaDrive Dataset summary TrueHuman PentaDrive is a compact behavioral matrix for affective and conversational modeling. It organizes human-relevant motivational structure into five drives (coded S, K, A, M, G), each implemented as a set of nodes (stable behavioral motifs). Every node defines three phases—anticipation, release, and block—with: Markers: short textual cues for lightweight classification or retrieval Response kernels: structured hints for… See the full description on the dataset page: https://huggingface.co/datasets/datamarketinglabs/pentadrive-v1.texttext-classification1K<n<10K0 likes26 downloads6mo agoHugging Face22pensieves /PreferenceTravelPlanner PreferenceTravelPlanner Dataset PreferenceTravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints and preferences. For more details, see our paper. It is created by augmenting TravelPlanner (See paper for more details) with several common type of preferences under various preference paradigms. Introduction In PreferenceTravelPlanner, for a given query, language agents are expected to formulate a… See the full description on the dataset page: https://huggingface.co/datasets/pensieves/PreferenceTravelPlanner.tabulartext-generation1K<n<10K0 likes25 downloads7mo agoHugging Face23penfever /MM-MathInstruct-longest-20k-solutions-with-images MM-MathInstruct Longest 20K Solutions with Images This dataset contains the top 20,000 samples from MathLLMs/MM-MathInstruct selected by solution length, filtered to include only samples with valid images. Dataset Structure Usage Source This dataset is derived from MathLLMs/MM-MathInstruct by selecting the 20,000 samples with the longest solution text. License Apache 2.0 (inherited from source dataset) textvisual-question-answering10K<n<100K0 likes22 downloads1y agoHugging Face24louisbrulenaudet /code-disciplinaire-penal-marine-marchande Code disciplinaire et pénal de la marine marchande, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-disciplinaire-penal-marine-marchande.tabulartext-generationn<1K0 likes20 downloads1y agoHugging Face25louisbrulenaudet /code-penitentiaire Code pénitentiaire, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-penitentiaire.tabulartext-generation1K<n<10K0 likes20 downloads2y agoHugging Face26louisbrulenaudet /code-pensions-retraite-marins-francais-commerce-peche-plaisance Code des pensions de retraite des marins français du commerce, de pêche ou de plaisance, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-pensions-retraite-marins-francais-commerce-peche-plaisance.tabulartext-generationn<1K0 likes18 downloads2y agoHugging Face27pengxiang /nap-parallel-packing-demo NAP Parallel Packing Demo Parallel-packed pretraining data built from FineWeb sample-10BT. Core idea: blocks within each sample are semantically related but not duplicates; block order is shuffled to break privileged sequential ordering. Format Each line in train.jsonl is a JSON object: { "text": "<blk>block 1 text</blk><blk>block 2 text</blk><blk>block 3 text</blk>", "blocks": ["block 1 text", "block 2 text", "block 3 text"], "metadata": {… See the full description on the dataset page: https://huggingface.co/datasets/pengxiang/nap-parallel-packing-demo.texttext-generation1K<n<10K1 likes17 downloads6mo agoHugging Face28pennydoesdev /Orb-training-data Orb Training Data Training dataset for Orb, an advanced AI coding and deployment assistant. Overview Examples: 201 chat conversations Format: JSONL with system/user/assistant messages Topics: Python, JavaScript/TypeScript, Go, Rust, DevOps, databases, security, deployment, testing, monitoring Orb's 12-Phase Workflow Each example teaches Orb to follow its structured workflow: Code review with pros/cons Debug guidance Deployment strategy Iterative debugging (5… See the full description on the dataset page: https://huggingface.co/datasets/pennydoesdev/Orb-training-data.texttext-generationn<1K0 likes16 downloads6mo agoHugging Face29louisbrulenaudet /code-pensions-militaires-invalidite-victimes-guerre Code des pensions militaires d'invalidité et des victimes de guerre, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-pensions-militaires-invalidite-victimes-guerre.tabulartext-generation1K<n<10K0 likes14 downloads2y agoHugging Face30Yansa /Pengaruhinternet Hermes Function-Calling V1 This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models. This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/Yansa/Pengaruhinternet.texttext-generation10K<n<100K0 likes14 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.