datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
massive-templates
Purpose. This dataset was collected specifically for intent-parser benchmarking, independently from any OVOS skill. Skill-derived utterances tend to overfit the exact phrasings a plugin was tuned on; this data is drawn from a disjoint source so it measures whether an OVOS intent plugin generalizes rather than memorizes. It is part of the OVOS intent-classification datasets used by the OVOS Plugin Arena intent benchmark.
Funding
Developed by TigreGotico for OpenVoiceOS as part… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/massive-templates.workflow-templatesovos-intents-v5-templates
OVOS intent corpus v6.1
This is the intent corpus for the OpenVoiceOS skill fleet. Each row is one utterance with the
intent label it belongs to. The corpus is built from the locale resource files of the skills
themselves, so it grows when the fleet gains locales.
Both splits ship here. Pin the tag v6.1 to get one build of both. The repository name says
v5 on purpose: consumers pin it by name, and the version of the content is the tag.
This is a candidate. The intent-engine lane… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intents-v5-templates.hass-intent-templates
Purpose. This dataset was collected specifically for intent-parser benchmarking, independently from any OVOS skill. Skill-derived utterances tend to overfit the exact phrasings a plugin was tuned on; this data is drawn from a disjoint source so it measures whether an OVOS intent plugin generalizes rather than memorizes. It is part of the OVOS intent-classification datasets used by the OVOS Plugin Arena intent benchmark.
Funding
Developed by TigreGotico for OpenVoiceOS as part… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/hass-intent-templates.legal-draft-templatesernest-nuclei-templates-v3
Ernest Nuclei Templates Dataset v3
A comprehensive dataset for training models to generate Nuclei security scanning templates from vulnerability descriptions.
Dataset Description
This dataset contains 11,590 training examples for generating Nuclei YAML templates in JSON IR (Intermediate Representation) format from structured vulnerability specifications.
Format
Each training example consists of:
id: Unique identifier (CVE ID, CWE ID, or template name)
prompt:… See the full description on the dataset page: https://huggingface.co/datasets/OzLabs/ernest-nuclei-templates-v3.ernest-nuclei-templates-v1-instruct
Ernest Nuclei Templates Dataset v1 (Instruction Format)
This is a restructured version of the Ernest Nuclei Templates dataset optimized for instruction tuning.
Format
Each example has this structure:
{
"instruction": "Generate a Nuclei template for CVE-2024-1234",
"input": "Title: XSS in Product\nSummary: Description...\nSeverity: high",
"output": "id: cve-2024-1234\n\ninfo:\n name: ...\n\nhttp:\n - method: GET\n ...",
"category": "cve",
"severity": "high"… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/ernest-nuclei-templates-v1-instruct.fullstack-templatesllm-rag-optimized-schema-templates
Schema.org JSON-LD Templates Optimized for LLM RAG Retrieval (2026)
Curated dataset of Schema.org JSON-LD templates designed, tested, and optimized for Retrieval-Augmented Generation (RAG) systems, SearchGPT, Gemini, and Claude search parsers.
Published by Pixel Office EU.
Purpose
Standard Schema.org markup is often too nested or dense for token-efficient LLM context window ingestion. These templates prioritize high-salience fields that crawlers prioritize when… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-rag-optimized-schema-templates.camera-shot-templates
Camera Shot Templates
镜头语言模板库:镜头类型、运镜、灯光、情绪、Kling 示例 Prompt。
professional-ai-file-naming-templates
📂 Renomee AI: 专业语义化文件重命名规则集
🚀 让文件系统具备“语义理解”能力
传统的批量重命名工具仅支持正则替换(Regex),无法理解文件内容的真实含义。Renomee AI 通过大语言模型(LLM)提取文档深层元数据,将混乱的原始文件名转化为具备高度可读性的结构化资产。
本项目开源了一套针对不同行业(金融、法律、学术、个人效率)的语义命名逻辑模版,旨在为 AI 自动化办公提供标准化参考。
💡 为什么需要语义重命名?
在处理海量文件时,人工命名的成本为 $O(n)$。而传统工具无法处理如下场景:
简历解析: 从 4686_模板.docx 中提取姓名和学校。
合同管理: 从 技术开发合同.pdf 中识别出甲方、乙方和项目周期。
财务审计: 自动从电子发票中抓取金额和日期。
📊 典型转换案例 (Case Studies)
场景
原始文件名 (Unstructured)
Renomee AI 重构后 (Structured)
简历招聘… See the full description on the dataset page: https://huggingface.co/datasets/tianhe2023/professional-ai-file-naming-templates.prompt-templates
Prompt Templates Dataset
Overview
This dataset contains 300 unique prompt templates designed for use with large language models (LLMs). Each prompt is structured as a JSON object, making it easy to integrate into machine learning pipelines, especially those using the HuggingFace ecosystem.
Features
300 unique prompts across various categories and difficulty levels.
Structured JSONL format for easy ingestion into datasets.
Diverse categories including creative… See the full description on the dataset page: https://huggingface.co/datasets/Kenshiii/prompt-templates.viral-structure-templates
Viral Structure Templates
爆款视频结构模板库:结构名、适用产品、Hook 风格、时长、剧情节拍。
fm_templatescf-myframes-templatesquestion_templates
