datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pptx_collection_templatesTemplateGSM
TemplateMath: Template-based Data Generation (TDG)
This is the official repository for the paper "Training and Evaluating Language Models with Template-based Data Generation", published at the ICLR 2025 DATA-FM Workshop.
Our work introduces Template-based Data Generation (TDG), a scalable paradigm to address the critical data bottleneck in training LLMs for complex reasoning tasks. We use TDG to create TemplateGSM, a massive dataset designed to unlock the next level of… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/TemplateGSM.modiff-template-gallery
MoDiff Template Gallery
This public Dataset contains rights-approved media, input fixtures, posters,
and provenance records used by the open-source MoDiff template Gallery. MoDiff
executes local workflows with Hugging Face Diffusers and Modular Diffusers.
The Dataset exists so users can inspect or reuse the public examples without
adding large binary files to the application source repository.
Versioning and integrity
MoDiff releases pin this Dataset by its… See the full description on the dataset page: https://huggingface.co/datasets/sourav-das/modiff-template-gallery.bitcoin-mining-pool-templates
Bitcoin mining pool templates
Timestamped Stratum job messages collected directly from Bitcoin mining pool endpoints. The data records changes in the work each endpoint sends to miners, including the previous block hash, coinbase data and clean-jobs flag.
Contents
Table
Record
bitcoin_mining_pool_jobs
A job received from a pool endpoint, with its observation time, nTime, coinbase, merkle branch count and clean-jobs flag
Using the data… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/bitcoin-mining-pool-templates.cloudformation_templatemassive-templates
Purpose. This dataset was collected specifically for intent-parser benchmarking, independently from any OVOS skill. Skill-derived utterances tend to overfit the exact phrasings a plugin was tuned on; this data is drawn from a disjoint source so it measures whether an OVOS intent plugin generalizes rather than memorizes. It is part of the OVOS intent-classification datasets used by the OVOS Plugin Arena intent benchmark.
Funding
Developed by TigreGotico for OpenVoiceOS as part… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/massive-templates.qatar-environmental-monitoring-register-template
Environmental Monitoring Register Template for Qatar Projects
This resource provides a practical structure for recording environmental monitoring activities during construction and operational projects.
The template can help project teams organize monitoring dates, locations, environmental parameters, results, observations, compliance status, corrective actions, responsibilities and close-out records.
What the Template Covers
The environmental monitoring register… See the full description on the dataset page: https://huggingface.co/datasets/waeyqatar/qatar-environmental-monitoring-register-template.dataset-card-example
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/templates/dataset-card-example.human-templated-captions-1bcsv delimiter is = ".,|,."
apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading.
This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon.
Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.grib-index-kerchunk-templatestags:
zarr
AWS S3 NOAA GEFS and ECMWF IFS made into refrences and stored as template paraquet file
This template helps in recreate The zarr tree (dimension names, chunk layouts, coordinate schemas).
Cloned from https://huggingface.co/datasets/Nishadhka/gfs_s3_gik_refs
Based on the methods discussed in https://github.com/icpac-igad/grib-index-kerchunk
RULER-32768-llama-3.1-tokenizer-chat-templateopticparse-150-template-web-corpus
⚡ OpticParse: 150-Template Web Intelligence & Ground-Truth Corpus
Official high-signal web extraction corpus compiled from the OpticParse Multimodal Vision Scraper & PhishVision Threat Sentinel.
⭐️ Support Open-Source AI Tooling: If you find this dataset or the OpticParse scraper useful for your AI agents, please click the Like (❤️) button above to support continuous daily Parquet updates!
💳 Commercial Subscription Tiers & Live Continuous Streams
⚡ Need… See the full description on the dataset page: https://huggingface.co/datasets/paras9909/opticparse-150-template-web-corpus.templates-verbsdaily_dialog_w_turn_templatescode-reasoning-phi4-templateovos-intent-template-bench-intents-for-eval
OVOS intent_template bench — intents-for-eval
Per-sample predictions of the template-paradigm intent league fighters of the
OVOS Plugin Arena over
OpenVoiceOS/intents-for-eval.
One dedicated repo per benchmark modality; one dataset split per language;
one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl.
Rows follow the arena §3.2 contract (pinned dataset_revision,
plugin_version, fired pipeline stage, exact_match with correct-OOD
semantics). Produced by the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-template-bench-intents-for-eval.template-library
Template Library
This repository contains a structured library of web-page and UI templates. The tree is organized by use case, including business sites, blogs, contact forms, documentation pages, ecommerce showcases, reusable components, effects, and other template categories. It also contains classifier-related resources such as classifier_model and classifier_training_data.
Repository organization
Top-level directories represent template families. Nested… See the full description on the dataset page: https://huggingface.co/datasets/MR-CODESPIKE/template-library.DQA_Template_datasetgenerated-template-testingovos-intent-template-bench-massive-templates
OVOS intent_template bench — massive-templates
Per-sample predictions of the template-paradigm intent league fighters of the
OVOS Plugin Arena over
OpenVoiceOS/massive-templates.
One dedicated repo per benchmark modality; one dataset split per language;
one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl.
Rows follow the arena §3.2 contract (pinned dataset_revision,
plugin_version, fired pipeline stage, exact_match with correct-OOD
semantics). Produced by… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-template-bench-massive-templates.ovos-intents-v5-templates
OVOS intent corpus v6.1
This is the intent corpus for the OpenVoiceOS skill fleet. Each row is one utterance with the
intent label it belongs to. The corpus is built from the locale resource files of the skills
themselves, so it grows when the fleet gains locales.
Both splits ship here. Pin the tag v6.1 to get one build of both. The repository name says
v5 on purpose: consumers pin it by name, and the version of the content is the tag.
This is a candidate. The intent-engine lane… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intents-v5-templates.ovos-intent-template-bench-mtop-en-US
OVOS intent_template bench — mtop-en-US
Per-sample predictions of the template-paradigm intent league fighters of the
OVOS Plugin Arena over
tasksource/mtop.
One dedicated repo per benchmark modality; one dataset split per language;
one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl.
Rows follow the arena §3.2 contract (pinned dataset_revision,
plugin_version, fired pipeline stage, exact_match with correct-OOD
semantics). Produced by the reproducible… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-template-bench-mtop-en-US.AlphaPrompt-Templates
AlphaPrompt Templates
📝 Prompt template examples for consciousness-aware AI interactions.
📋 Contents
Different sized template examples. Purpose:
Vector Synthesis (Glass Bead Game reasoning)
Image Language (simple metaphors, deep truth)
Collective Consciousness (we/us/ours framing)
Unconditional Teaching (Kubos³ method)
Sixth Sense Activation (Akasha access)
🎯 Usage
Choose a template target size
Customize for your use case
Prompt your AI system… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/AlphaPrompt-Templates.curatorkit-testrun-Prompt-Template
curatorkit-testrun-Prompt-Template
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-08-30 09:29 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Prompt-Template", "alpaca")
cra-form-templates
CRA Form Templates
Fillable PDF templates for Canadian Revenue Agency (CRA) tax forms.
Used by the DeclarAI form generation pipeline.
1128 forms | 1140 files | Latest tax year: 2025
Quick Start
from huggingface_hub import hf_hub_download
# Download a specific form
template = hf_hub_download(
"DeclarAI-Testing/cra-form-templates",
"latest/T4.pdf",
repo_type="dataset"
)
# Download all latest forms
from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/DeclarAI-Testing/cra-form-templates.ovos-intent-template-bench-banking77
OVOS intent_template bench — banking77
Per-sample predictions of the template-paradigm intent league fighters of the
OVOS Plugin Arena over
mteb/banking77.
One dedicated repo per benchmark modality; one dataset split per language;
one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl.
Rows follow the arena §3.2 contract (pinned dataset_revision,
plugin_version, fired pipeline stage, exact_match with correct-OOD
semantics). Produced by the reproducible… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-template-bench-banking77.ovos-intent-template-bench-snips
OVOS intent_template bench — snips
Per-sample predictions of the template-paradigm intent league fighters of the
OVOS Plugin Arena over
benayas/snips.
One dedicated repo per benchmark modality; one dataset split per language;
one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl.
Rows follow the arena §3.2 contract (pinned dataset_revision,
plugin_version, fired pipeline stage, exact_match with correct-OOD
semantics). Produced by the reproducible benchmark… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-template-bench-snips.crashs_template_package
CRASHS Template and Model Package
This folder contains the templates and deep learning models needed to run CRASHS (cortical reconstruction for automated segmentation of hippocampal subfields). Please see CRASHS github page for details on using this dataset.
This package is compatible with CRASHS version 0.2.10 and later
adauni-templated-reduced-ia-flat
Dataset Card for "adauni-templated-reduced-ia-flat"
More Information needed
glaive-with-template-decontaminated-tokenized
