datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ernest-nuclei-templates-v3
Ernest Nuclei Templates Dataset v3
A comprehensive dataset for training models to generate Nuclei security scanning templates from vulnerability descriptions.
Dataset Description
This dataset contains 11,590 training examples for generating Nuclei YAML templates in JSON IR (Intermediate Representation) format from structured vulnerability specifications.
Format
Each training example consists of:
id: Unique identifier (CVE ID, CWE ID, or template name)
prompt:… See the full description on the dataset page: https://huggingface.co/datasets/OzLabs/ernest-nuclei-templates-v3.dfm11-danish-template-instantiator-training
dfm11-danish-template-instantiator-training
Independently accepted Danish template-instantiator supervision produced by the DFM-owned FineInstructions reproduction pipeline.
Rows retain generation and audit provenance. Local filesystem paths are removed.
The synthetic release does not broaden rights attached to upstream grounding
or query sources; consult each row's source provenance and upstream terms.
nuclei-template-generation-dataset-2.3K
nuclei-template-generation-dataset-2.3K
Description:
A specialized instruction-tuning dataset of 2350 examples for training large language models to generate Nuclei YAML templates. Each example consists of a fixed instruction, a structured JSON input describing a vulnerability (CVE, product, HTTP details, detection logic), and the corresponding valid Nuclei template as output. The dataset was constructed from the official Nuclei Templates repository (HTTP… See the full description on the dataset page: https://huggingface.co/datasets/NormanRey/nuclei-template-generation-dataset-2.3K.llm-rag-optimized-schema-templates
Schema.org JSON-LD Templates Optimized for LLM RAG Retrieval (2026)
Curated dataset of Schema.org JSON-LD templates designed, tested, and optimized for Retrieval-Augmented Generation (RAG) systems, SearchGPT, Gemini, and Claude search parsers.
Published by Pixel Office EU.
Purpose
Standard Schema.org markup is often too nested or dense for token-efficient LLM context window ingestion. These templates prioritize high-salience fields that crawlers prioritize when… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-rag-optimized-schema-templates.professional-ai-file-naming-templates
📂 Renomee AI: 专业语义化文件重命名规则集
🚀 让文件系统具备“语义理解”能力
传统的批量重命名工具仅支持正则替换(Regex),无法理解文件内容的真实含义。Renomee AI 通过大语言模型(LLM)提取文档深层元数据,将混乱的原始文件名转化为具备高度可读性的结构化资产。
本项目开源了一套针对不同行业(金融、法律、学术、个人效率)的语义命名逻辑模版,旨在为 AI 自动化办公提供标准化参考。
💡 为什么需要语义重命名?
在处理海量文件时,人工命名的成本为 $O(n)$。而传统工具无法处理如下场景:
简历解析: 从 4686_模板.docx 中提取姓名和学校。
合同管理: 从 技术开发合同.pdf 中识别出甲方、乙方和项目周期。
财务审计: 自动从电子发票中抓取金额和日期。
📊 典型转换案例 (Case Studies)
场景
原始文件名 (Unstructured)
Renomee AI 重构后 (Structured)
简历招聘… See the full description on the dataset page: https://huggingface.co/datasets/tianhe2023/professional-ai-file-naming-templates.
