instruction-training
constrained-instruction-training-pool
Constrained instruction training pool
Public prompts for writing tasks, many of them carrying a constraint a program can check, from ten
datasets read at the pinned revisions named below and one layer built here from them. The pool is
laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 462652 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/constrained-instruction-training-pool.instruction-following-training-pool
Instruction following training pool
Public prompts for writing tasks, many of them carrying a constraint a program can check, from
eight datasets read at the pinned revisions named below and one layer built here from them. The
pool is laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 310602 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/instruction-following-training-pool.System-Prompt-Instruction-Real-world-Implementation-Training-set
SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set)
Dataset Summary
SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/EricLu/System-Prompt-Instruction-Real-world-Implementation-Training-set.PartC_Training_Instruction_Model_V1_DatasetMega-Instruction-Following-Dataset-Queso-Training
