datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TemplateGSM
TemplateMath: Template-based Data Generation (TDG)
This is the official repository for the paper "Training and Evaluating Language Models with Template-based Data Generation", published at the ICLR 2025 DATA-FM Workshop.
Our work introduces Template-based Data Generation (TDG), a scalable paradigm to address the critical data bottleneck in training LLMs for complex reasoning tasks. We use TDG to create TemplateGSM, a massive dataset designed to unlock the next level of… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/TemplateGSM.curatorkit-testrun-Prompt-Template
curatorkit-testrun-Prompt-Template
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-08-30 09:29 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Prompt-Template", "alpaca")
star-dataset-templates
STAR Templates
STAR Templates is a curated collection of 355 Jinja2 instruction templates for Arabic NLP tasks, spanning 27 tasks across 87 source datasets, contributed by 7 prompters. The templates were authored collaboratively on PromptLab. This dataset and the experiments built on it are described in STAR: instruction tuning for Arabic across tasks, datasets, and models.
For these templates rendered against datasets samples, see the companion dataset: STAR Instructions.
📦… See the full description on the dataset page: https://huggingface.co/datasets/KFUPM-JRCAI/star-dataset-templates.legal-templates-multilingual
Forms Legal — Multilingual Legal Templates Corpus
19,419 legal document templates across 36 jurisdictions in 12 languages, expanded to 22,876 (document × locale) rows. Released under CC-BY-4.0 by forms-legal.com.
Quick description
A multilingual corpus of structured legal document templates spanning 36 jurisdictions. Each document includes a multi-section editorial brief (whatIs, whenNeeded, keyElements, howToFill, legalRequirements, commonMistakes), 5-8… See the full description on the dataset page: https://huggingface.co/datasets/forms-legal/legal-templates-multilingual.Primus-Reasoning-DeepSeek-Qwen-Template
Primus-Reasoning-DeepSeek-Qwen-Template
This dataset is a converted version of trendmicro-ailab/Primus-Reasoning
adapted for DeepSeek-Qwen template format.
Changes
The original dataset used custom special tokens for reasoning:
<|reserved_special_token_0|>{reasoning}<|reserved_special_token_1|>{answer}
This version has been converted to use DeepSeek-Qwen's think tags:
<think>{reasoning}</think>{answer}
Dataset Structure
Each example contains:
prompt: The… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Primus-Reasoning-DeepSeek-Qwen-Template.
