qrk-labs/akeel-thought-injection-50k-augmented
Akeel Thought Injection Dataset (50K Augmented) Multi-turn augmented reasoning traces for robust thought injection training A QRK Labs Research Dataset Overview This dataset contains 50,000 augmented examples designed to prevent overfitting when training thought injection models. It expands the base 20K dataset through: Shuffled originals (17,500) — Base samples randomly reordered System prompt variations (17,500) — Same Q&A with… See the full description on the dataset page: https://huggingface.co/datasets/qrk-labs/akeel-thought-injection-50k-augmented.
<div align="center">
<!-- Muqarnas Layers Logo --> <svg width="80" height="80" viewBox="0 0 64 64" fill="none" xmlns="http://www.w3.org/2000/svg"> <path d="M32 4 L4 32 L32 60 L60 32 Z" fill="none" stroke="#fafafa" stroke-width="0.6" opacity="0.3"/> <path d="M32 10 L10 32 L32 54 L54 32 Z" fill="none" stroke="#fafafa" stroke-width="0.8" opacity="0.5"/> <path d="M32 16 L16 32 L32 48 L48 32 Z" fill="none" stroke="#fafafa" stroke-width="1" opacity="0.7"/> <path d="M32 22 L22 32 L32 42 L42 32 Z" fill="none" stroke="#fafafa" stroke-width="1.2"/> <circle cx="32" cy="32" r="6" fill="none" stroke="#fafafa" stroke-width="0.8"/> <circle cx="32" cy="32" r="2.5" fill="#fafafa"/> </svg>
Akeel Thought Injection Dataset (50K Augmented)
Multi-turn augmented reasoning traces for robust thought injection training
A QRK Labs Research Dataset
</div>
Overview
This dataset contains 50,000 augmented examples designed to prevent overfitting when training thought injection models. It expands the base 20K dataset through:
- Shuffled originals (17,500) — Base samples randomly reordered
- System prompt variations (17,500) — Same Q&A with different instruction phrasings
- Multi-turn conversations (15,000) — 2-4 chained Q&A pairs per conversation
Why Augmentation?
Single-turn training can cause models to overfit to specific response patterns. This augmented dataset helps models:
- Maintain context across multiple turns
- Generalize across different system prompts
- Learn robust thought injection behavior
Format
Single-turn (same as base dataset)
<|im_start|>system
You are a knowledgeable assistant...<|im_end|>
<|im_start|>user
Question here<|im_end|>
<|im_start|>assistant
<think>...</think>
Answer here<|im_end|>Multi-turn (2-4 exchanges)
<|im_start|>system
You are an AI assistant with access to external knowledge...<|im_end|>
<|im_start|>user
First question<|im_end|>
<|im_start|>assistant
<think>...</think>
First answer<|im_end|>
<|im_start|>user
Second question<|im_end|>
<|im_start|>assistant
<think>...</think>
Second answer<|im_end|>Statistics
Augmentation Details
System Prompt Variations
Five equivalent phrasings are used:
- "You are a knowledgeable assistant that provides accurate, helpful answers..."
- "You are an AI assistant with access to external knowledge..."
- "You are a helpful assistant. When answering questions, use
<knowledge>tags..." - "You are an intelligent assistant that retrieves information..."
- "You are a knowledgeable AI. Request external information with
<knowledge>tags..."
Multi-turn Construction
- Randomly samples 2-4 independent Q&A pairs
- Chains them into a single conversation
- Assigns a random system prompt variation
- Final shuffle ensures no ordering bias
Usage
from datasets import load_dataset
dataset = load_dataset("qrk-labs/akeel-thought-injection-50k-augmented")Related
- [akeel-thought-injection-20k](https://huggingface.co/datasets/qrk-labs/akeel-thought-injection-20k) — Base single-turn dataset
- [akeel-cot](https://huggingface.co/qrk-labs/akeel-cot) — Research model trained on this data
License
Apache 2.0
<div align="center"> <sub>Built with ☁️ by QRK Labs</sub> </div>
