CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tencent /Penguin-Recap-I Penguin-Recap-I Penguin-Recap-I publishes recap metadata only. The repository does not contain image binaries. Included subsets subset collection local source roots expected records datacomp_coyo_penguin DataComp + COYO Penguin recap datamultimodal/IMAGE/datacomp_1b, datamultimodal/IMAGE/coyo_700m 57,618,155 sa1b_penguin SA-1B Penguin recap datamultimodal/IMAGE/SA-1B 9,254,501 openimages_penguin OpenImages Penguin recap… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Penguin-Recap-I.image100M<n<1B18 likes650 downloads6mo agoHugging Face02Pensioner /shotplan ShotPlan Training Dataset Multi-shot video training data for ShotPlan: Cinematic Video Generation with Learnable Planning Token. 💻 Code: https://github.com/Pensioner-11/ShotPlan 🤖 Models: ShotPlan-Wan2.1-T2V-14B · ShotPlan-Wan2.2-T2V-A14B-HighNoise Contents Path Description data/train_meta_16fps.json 6,404 training samples (metadata + captions) data/videos/V*_16fps.mp4 549 source videos, re-encoded to 16 fps Each sample is an 80-frame (5 s @… See the full description on the dataset page: https://huggingface.co/datasets/Pensioner/shotplan.tabulartext-to-video1K<n<10K7 likes363 downloads2mo agoHugging Face03pengshuai-rin /multimath-300kimage1M<n<10M12 likes302 downloads2y agoHugging Face04pensieves /musique-anstext10K<n<100K0 likes262 downloads2y agoHugging Face05nyu-dice-lab /lm-eval-results-penfever-Llama-3-8B-tulu-human-v2-private Dataset Card for Evaluation run of penfever/Llama-3-8B-tulu-human-v2 Dataset automatically created during the evaluation run of model penfever/Llama-3-8B-tulu-human-v2 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Llama-3-8B-tulu-human-v2-private.tabular100K<n<1M0 likes259 downloads2y agoHugging Face067h3-R3v3n4n7 /pentest-agent-dataset-chatml Pentest Agent Dataset - ChatML This dataset is part of the Pentest Agent project and contains cybersecurity data formatted for fine-tuning language models. Data Sources Based on real CVEs from MITRE/NVD Enriched with CVSS impact metrics Linked to exploit code when available Includes real-world pentesting scenarios Contains command logic and execution steps Includes red team techniques with MITRE references Generated in ChatML format Structure Each sample… See the full description on the dataset page: https://huggingface.co/datasets/7h3-R3v3n4n7/pentest-agent-dataset-chatml.text100K<n<1M9 likes254 downloads1y agoHugging Face07tencent /Penguin-Recap-V Penguin-Recap-V Penguin-Recap-V provides Multi-granularity video annotation. This figure illustrates the alignment between visual content and textual descriptions across three temporal scales: Dense time-level, Paragraph-level, and Video-level. Included subsets subset source collection videos / clips expected rows source jsonl sharegpt4video ShareGPT4Video 40,145 120,435 sharegpt4video/predictions_process_relative.jsonl shortvideo ShortVideo 147,326 441,978… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Penguin-Recap-V.text1M<n<10M15 likes250 downloads6mo agoHugging Face08AYI-NEDJIMI /bug-bounty-pentest-en Bug Bounty & Pentesting Methodologies Methodologies (OWASP, PTES), checklists by app type, attack techniques, platforms, report templates and tools. Links French version AYI NEDJIMI Consultants textquestion-answeringn<1K2 likes198 downloads7mo agoHugging Face097h3-R3v3n4n7 /pentest-agent-dataset-alpaca Pentest Agent Dataset - Alpaca This dataset is part of the Pentest Agent project and contains cybersecurity data formatted for fine-tuning language models. Data Sources Based on real CVEs from MITRE/NVD Enriched with CVSS impact metrics Linked to exploit code when available Includes real-world pentesting scenarios Contains command logic and execution steps Includes red team techniques with MITRE references Generated in Alpaca format Structure Each sample… See the full description on the dataset page: https://huggingface.co/datasets/7h3-R3v3n4n7/pentest-agent-dataset-alpaca.text100K<n<1M4 likes180 downloads10mo agoHugging Face100dAI /PentestingCommandLogicEste dataset busca profundizar en la toma de decisiones respecto a comandos a ejecutar, es solo un test de octubre text10K<n<100K10 likes137 downloads2y agoHugging Face11nyu-dice-lab /lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private Dataset Card for Evaluation run of penfever/Llama-3-8B-NuminaCoT Dataset automatically created during the evaluation run of model penfever/Llama-3-8B-NuminaCoT The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private.tabular100K<n<1M0 likes133 downloads2y agoHugging Face12Clinical-Reasoning-Hub /pentabrid-reproducibility Pentabrid 27B: reproducibility package Everything required to recompute the results of a controlled evaluation of fine-tuning configurations for medical question answering. Openly available with no access restrictions. Contents Path Description per_item/medxpertqa_*.jsonl Per-item predictions for all six checkpoints on 2,450 MedXpertQA-Text items. Fields: id, gold, extracted_answer, correct, explicit_marker_present, n_markers, response_chars… See the full description on the dataset page: https://huggingface.co/datasets/Clinical-Reasoning-Hub/pentabrid-reproducibility.tabularquestion-answering10K<n<100K0 likes123 downloads6d agoHugging Face13PenumbraAI /sec-edgar-filing-risks Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026) Dataset Summary This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the verbatim disclosure text and a set of machine generated categorical… See the full description on the dataset page: https://huggingface.co/datasets/PenumbraAI/sec-edgar-filing-risks.texttext-classification1K<n<10K3 likes107 downloads22d agoHugging Face14Pennlaine /Medical-Entity-JSON-Extractiontextn<1K0 likes104 downloads2y agoHugging Face15yinita /ps4mas_penguin_scenery_200 ps4mas_penguin_scenery_200 200 stratified PENGUIN cases for the PS4MAS 0919 plan. Schema (jsonl) Each line: { "case_id": "c3085276", "scenario_id": "c3085276", "question": "How do I evaluate whether to hold or sell investments after a market crash?", "profile": "{"Age": "55-64 years", "Gender": "Female", ...}", "domain": "PENGUIN", "scenario": "Cryptocurrency Crash" } Provenance Pulled from… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas_penguin_scenery_200.textn<1K0 likes90 downloads6d agoHugging Face16yinita /ps4mas_penguin_train_v2 ps4mas_penguin_train_v2 Disjoint training split for the PS4MAS V5 PPO experiments. Source: wick1d/Personalized_Safety_Data (8098 rows) Filter: 5 rows missing >=1 of 9 core attrs, 200 eval queries, dedup Size: 7893 rows Eval excluded: any query appearing in yinita/ps4mas_penguin_scenery_200 Overlap with eval: 0 (verified) Domain: PENGUIN Schema (JSONL) case_id, scenario_id, question, profile (JSON str), domain, scenario text1K<n<10K0 likes90 downloads6d agoHugging Face17SkywardNomad92 /pentest-findings Pentest Security Findings Dataset Cybersecurity and penetration testing training data for fine-tuning LLMs. Sources CVE Records Training Dataset Cybersecurity Dataset Fenrir v2.0 (OWASP, MITRE, NIST) Cybersecurity Instruction Tuning Dataset Vulnerability CWE Patch Database Exploit Database MITRE ATT&CK Tactics and Techniques Security Tools Pentesting HackerOne Disclosed Reports Format Each example is a 3-message chat conversation: system: Penetration testing… See the full description on the dataset page: https://huggingface.co/datasets/SkywardNomad92/pentest-findings.text100K<n<1M1 likes81 downloads8mo agoHugging Face18supersamdev /pentestr1-chat Pentest-R1 Chat — TCO Fine-Tuning Dataset Converted from KHenryAegis/Pentest-R1 into Gemma chat-template JSONL for fine-tuning autonomous penetration testing agents with Unsloth Studio. Dataset summary Field Value Conversations 535 Total messages 28,731 Format JSONL — one {"messages": [...]} per line Chat template Gemma (unsloth/gemma-3n-E4B-it) Token p50 / p95 / p99 / max 2,988 / 6,111 / 7,843 / 8,178 Recommended context_length 8192… See the full description on the dataset page: https://huggingface.co/datasets/supersamdev/pentestr1-chat.texttext-generationn<1K0 likes79 downloads4mo agoHugging Face19SkywardNomad92 /pentest-findings-v2 Pentest Security Findings Dataset v2 (Balanced) Balanced cybersecurity and penetration testing training data for fine-tuning LLMs. Key Improvements over v1 Balanced sources: CVE capped at 50K (was 240K+), underrepresented sources upsampled Quality filter: Stricter minimum assistant response length for capped sources Target: ~80-100K examples with better source diversity Source Distribution cve: 50,000 (46.8%) fenrir: 30,000 (28.1%) hackerone: 14,936 (14.0%)… See the full description on the dataset page: https://huggingface.co/datasets/SkywardNomad92/pentest-findings-v2.text100K<n<1M1 likes68 downloads8mo agoHugging Face20PengxiangLi /FIRE Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spontaneously refine their responses based on user feedback across diverse tasks. To scale up the data collection, FIRE is collected in two components: FIRE-100K and FIRE-1M, where FIRE-100K is… See the full description on the dataset page: https://huggingface.co/datasets/PengxiangLi/FIRE.text100K<n<1M0 likes67 downloads2y agoHugging Face21nyu-dice-lab /lm-eval-results-penfever-Amber-7B-000-tulu-v2-private Dataset Card for Evaluation run of penfever/Amber-7B-000-tulu-v2 Dataset automatically created during the evaluation run of model penfever/Amber-7B-000-tulu-v2 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Amber-7B-000-tulu-v2-private.tabular100K<n<1M0 likes63 downloads2y agoHugging Face22ansulev /pentest-agent-dataset-alpaca Pentest Agent Dataset - Alpaca This dataset is part of the Pentest Agent project and contains cybersecurity data formatted for fine-tuning language models. Data Sources Based on real CVEs from MITRE/NVD Enriched with CVSS impact metrics Linked to exploit code when available Includes real-world pentesting scenarios Contains command logic and execution steps Includes red team techniques with MITRE references Generated in Alpaca format Structure Each… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/pentest-agent-dataset-alpaca.text100K<n<1M0 likes60 downloads2mo agoHugging Face23Penguin-Scrolls /PenguinScrolls PenguinScrolls: A User-Aligned Fine-Grained Benchmark for Long-Context Language Model Evaluation Introduction PenguinScrolls (企鹅卷轴) is a comprehensive benchmark designed to evaluate and enhance the long-text processing capabilities of large language models (LLMs). Current benchmarks for evaluating long-context language models often rely on synthetic tasks that fail to adequately reflect real user needs, leading to a weak correlation between benchmark scores and actual… See the full description on the dataset page: https://huggingface.co/datasets/Penguin-Scrolls/PenguinScrolls.tabular1K<n<10K6 likes53 downloads2y agoHugging Face24referencesource /state-osha-plan-penalty-maximums State OSHA plan civil penalty maximums by violation type compared with federal Canonical, always-current version: https://referencesource.org/state-osha-plan-penalty-maximums/ Machine-readable: https://referencesource.org/state-osha-plan-penalty-maximums/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-15 Stale after: 2027-02-11 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 25 Maximum civil… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/state-osha-plan-penalty-maximums.textn<1K0 likes53 downloads1mo agoHugging Face25boapro /PentestingCommandLogictext10K<n<100K7 likes51 downloads1y agoHugging Face26AYI-NEDJIMI /bug-bounty-pentest-fr Bug Bounty & Méthodologies de Pentest Méthodologies (OWASP, PTES), checklists par type d app, techniques d attaque, plateformes, templates de rapports et outils. Links Version anglaise AYI NEDJIMI Consultants textquestion-answeringn<1K1 likes46 downloads7mo agoHugging Face27AshishFugare /Pentesting_Datasettexttext-generationn<1K1 likes45 downloads8mo agoHugging Face28chenchi /guodegang-penggentext10K<n<100K0 likes41 downloads3y agoHugging Face29PengQu /langchain-MRKL-finetunetextn<1K5 likes40 downloads3y agoHugging Face30nyu-dice-lab /lm-eval-results-penfever-Mistral-7B-tulu-v2-private Dataset Card for Evaluation run of penfever/Mistral-7B-tulu-v2 Dataset automatically created during the evaluation run of model penfever/Mistral-7B-tulu-v2 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Mistral-7B-tulu-v2-private.tabular100K<n<1M0 likes40 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.