datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Pro-Agent.Ox-Alpha-10k
Ox Alpha - 10k
10,005 single-turn prompts for text-response teacher generation
Each row carries id, category, subcategory
All data was gathered using stealth/ox-alpha via OpenRouter (reasoning effort high)
Topic distribution
Category
Rows
Share
Coding (incl. Go/Rust, C++/Java/C#, shell/CLI)
944
9.5%
Knowledge QA
891
9.0%
Logical reasoning & decisions
734
7.4%
Web development
720
7.2%
Game development
720
7.2%
Three.js / browser 3D
620
6.2%… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Ox-Alpha-10k.Fable-5-Cursor-TracesFable 5 Cursor Traces
244 Fable 5 Cursor agent sessions for training & research.
This dataset has 244 Cursor sessions with Fable 5 at High/xHigh/Max effort levels for distillation.
[!IMPORTANT]
This dataset is compatible with Teich! Use it directly in your Teich training pipeline.
[!WARNING]
The longest rows exceed one million characters of content. Apply prepare_data() with your intended tokenizer and an explicit context/oversize policy before training.… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Fable-5-Cursor-Traces.claude-4.5-opus-high-reasoning-250xThis is a reasoning dataset created using Claude Opus 4.5 with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Claude Opus 4.5 by fine-tuning already existing open-source LLMs.
Stats
Costs: $ 52.3 (USD)
Total tokens (input + output): 2.13 M
teich-test-v1
hy3-preview coding agent traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by tencent/hy3-preview:free.
Training-ready tools
Use this tools payload when rendering converted examples through your training chat template.
The same structure is emitted on each converted example as the tools field.
[
{
"type": "function",
"function": {
"name": "bash",
"description": "Execute bash… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/teich-test-v1.lordx64-claude-opus-4.7-max-cleaned
reasoning-distill-claude-opus-4-7-max-cleaned
Cleaned version of lordx64/reasoning-distill-claude-opus-4-7-max.
See the original dataset for full provenance, collection methodology, and terms of use.
Cleaning steps
Step
Filter
Reason
Rows removed
1
Simulated thinking (...)
Rows with ... in thinking/response indicate the model learned to simulate reasoning (e.g., "Now I'm laying out the puzzle grids...") rather than actually performing it. This causes failures… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/lordx64-claude-opus-4.7-max-cleaned.DeepSeek-v4-Flash-ChatThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Teich Test
This directory contains newline-delimited JSON training examples generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-flash.
Rows: 6313
Format
Each file is newline-delimited JSON where every line is already a training example.
Chat-only datasets include messages… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Flash-Chat.gpt-5.2-high-reasoning-250x
Generated using DataGen by TeichAI
This is a reasoning dataset created using GPT 5.2 with a reasoning depth set to high.
The dataset is meant for creating distilled versions of GPT 5.2 by fine-tuning already existing open-source LLMs.
Stats
Costs: $ 10.58 (USD)
Total tokens (input + output): N\A
gemini-3-flash-preview
Gemini 3 Flash Preview
This is a reasoning dataset created using Gemini 3 Flash Preview with a reasoning depth set to high.
The dataset is meant for creating distilled versions of Gemini 3 Flash Preview by fine-tuning already existing open-source LLMs.
This dataset contains a collection of prompts categorized by themes such as benchmarks, psychology, web development, embedded systems, creative writing, design, finance, legal, marketing, and science. No real benchmark questions were… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/gemini-3-flash-preview.gpt-5.1-high-reasoning-1000xThis is a reasoning dataset created using GPT 5.1 (high reasoning) from OpenAI. Some of these questions are from reedmayhew and the rest were generated.
Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting.
The dataset is meant for creating distilled versions of GPT 5.1 (high reasoning) thinking by fine-tuning already existing open-source LLMs.
Stats
Costs: $ 36.4 (USD)
Tokens: 3.94 M (total)
convo-v1A conversation dataset generated using MiniMax M2.1 to help teach models how to talk to users.
TeichAI-thinking-reasoning-x
TeichAI Thinking & Reasoning Datasets
A collection of prompts answered by large language models (LLMs) such as Google Gemini and OpenAI ChatGPT, with long-form reasoning enabled.
These datasets were originally created by TeichAI for distillation and reasoning-focused training workflows.
Schema
Each row in the dataset has the following fields:
question_hash: Truncated, base64-encoded MD5 hash of the question, useful for filtering and deduplication.
question: The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/TeichAI-thinking-reasoning-x.claude-haiku-4.5-high-reasoning-1700xThis is a reasoning dataset created using Claude Haiku 4.5 with reasoning effort set to high.
The dataset is meant for creating distilled versions of Claude Haiku 4.5 by fine-tuning already existing open-source LLMs.
This dataset includes an addition to our recently enhanced set of prompts to cover creative writing and multilingual creative writing.
Stats
Costs: $ 33.52 (USD)
Total tokens (input + output): 6.79 M
gpt-5-codex-1000xThis is a reasoning dataset created using GPT 5 Codex from OpenAI. Some of these questions are from reedmayhew and the rest were generated.
Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting.
The dataset is meant for creating distilled versions of GPT 5 Codex by fine-tuning already existing open-source LLMs.
Pony-Alpha-15k
Pony Alpha 15k
This is a reasoning dataset generated using the stealth model Pony Alpha, which ended up being GLM-5.
As the largest dataset we have made yet. The prompts from this dataset were almost all generated by GPT 5.1 and Gemini 3 (flash and pro).
The categories covered include academia, multi-lingual creative writing, finance, health, law, marketing/SEO, programming, philosophy, web dev, python scripting, and science.
Stats:
Cost: $ 0 (USD)
Tokens (input + output): 43.3 M
Gemini-3-Flash-Preview-VIBE
Gemini 3 Flash Preview VIBE
This dataset is our first attempt at an agentic coding SFT dataset.
All of the prompts for this dataset were sourced from MiniMaxAI/VIBE.
Each prompt was given to Gemini 3 Flash Preview with the follow tools and system prompt:
read_file - Read file contents from workspace
write_file - Write content to a file
edit_file - Replace text in a file
list_directory - List files and directories
search_code - Search for patterns in files
run_command - Execute… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Gemini-3-Flash-Preview-VIBE.Claude-Opus-Dataclaw-Unredacted
Claude Opus Dataclaw Unredacted
How this dataset was built
Collected the local Petromallet raw export plus selected public Dataclaw uploads.
Filtered to the supported Opus-family source rows.
Deduplicated by session_id and first user message.
Converted raw assistant tool_uses directly into structured OpenAI-style tool_calls.
Derived per-row tool definitions from canonical schemas and observed tool usage.
Preserved assistant reasoning in <think>...</think> blocks.… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Claude-Opus-Dataclaw-Unredacted.gemini-3-pro-preview-high-reasoning-1000xThis is a reasoning dataset created using Gemini 3 Pro Preview with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Gemini 3 Pro Preview by fine-tuning already existing open-source LLMs on the summarized reasoning traces provided from their API.
This dataset includes 250x from TeichAI/gemini-3-pro-preview-high-reasoning-250x
Stats
Costs: $ 32.7 (USD)
Tokens: 2.73 M… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/gemini-3-pro-preview-high-reasoning-1000x.grok-code-fast-1-1000xThis is a reasoning dataset created using Grok Code Fast 1 from xAI. Some of these questions are from reedmayhew and the rest were generated.
Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting.
The dataset is meant for creating distilled versions of Grok Code Fast 1 by fine-tuning already existing open-source LLMs.
claude-sonnet-4.5-high-reasoning-250xThis is a reasoning dataset created using Claude Sonnet 4.5 with a high reasoning effort. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Claude Sonnet 4.5 by fine-tuning already existing open-source LLMs.
The default system prompt from OpenrouterAI was used
You are Claude Sonnet 4.5, a large language model from anthropic.
Formatting Rules:
- Use Markdown for lists, tables, and styling.
- Use ```code fences```… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/claude-sonnet-4.5-high-reasoning-250x.gemini-3-pro-preview-high-reasoning-250xThis is a reasoning dataset created using Gemini 3 Pro Preview with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Gemini 3 Pro Preview by fine-tuning already existing open-source LLMs on the summarized reasoning traces provided from their API.
Stats
Costs: $ 5.78
Total tokens (input + output): 484 K
MiniMax-M2.1-Code-SFT
MiniMax M2.1 Code SFT
200 of the prompts for this dataset were sourced from MiniMaxAI/VIBE. The rest were generated.
Each prompt was given to MiniMax M2.1 with the follow tools and system prompt:
read_file - Read file contents from workspace
write_file - Write content to a file
edit_file - Replace text in a file
list_directory - List files and directories
search_code - Search for patterns in files
run_command - Execute shell commands (with timeout)
web_search - Web search (powered… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/MiniMax-M2.1-Code-SFT.polaris-alpha-1000xThis is a non-reasoning dataset created using Polaris Alpha (likely an alpha version of a new ChatGPT) from OpenAI. Some of these questions are from reedmayhew and the rest were generated.
Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting.
The dataset is meant for creating distilled versions of Polaris Alpha by fine-tuning already existing open-source LLMs.
TeichAI-ClaudeOpus4.5-High-CLEANEDNote:
Base datasets: TeichAI/claude-4.5-opus-high-reasoning-250x
Cleaned: low-quality data removed and content condensed (8MB -> 5MB)
It cost me $2 (USD)
[just kidding]
mistral-small-creative-500x
Mistral Small Creative - 500x
This is a non-reasoning dataset created using Mistral Small Creative.
The dataset is meant for creating distilled versions of Mistral Small Creative by fine-tuning already existing open-source LLMs.
This dataset only covers generating short fictional stories.
Meant for testing purposes -> to evaluate if creating a larger dataset would be worth it
Stats
Costs: $ 0.20 (USD)
Total tokens (input + output): 704K
kimi-k2-thinking-1000xThis is a reasoning dataset created using Kimi k2 thinking from MoonshotAI. Some of these questions are from reedmayhew and the rest were generated.
Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting.
The dataset is meant for creating distilled versions of Kimi k2 thinking by fine-tuning already existing open-source LLMs.
gpt-5.1-codex-max-1000x
GPT 5.1 Codex Max - 1,000x
This is a reasoning dataset created using GPT 5.1 Codex Max with a reasoning depth set to high.
The dataset is meant for creating distilled versions of GPT 5.1 Codex Max by fine-tuning already existing open-source LLMs.
This datasets prompts are adapted versions of our standard set of prompts, designed to provoke complex reasoning.
Stats
Costs: $ 23.51 (USD)
Total tokens (input + output): 2.39 M
Generated using DataGen by TeichAI
Hunter-Alpha-16k
Hunter Alpha 16k
This is a reasoning dataset generated using the stealth model Hunter Alpha, which was revealed to be xiaomi/mimo-v2-pro.
As the largest dataset we have made yet. The prompts from this dataset were almost all generated by Opus 4.5/4.6, GPT 5.1 and Gemini 3 (flash and pro).
The categories covered include academia, multi-lingual creative writing, finance, health, law, marketing/SEO, programming, philosophy, web dev, python scripting, and science.
Stats:
Cost: $ 0… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Hunter-Alpha-16k.deepseek-v3.2-speciale-openr1-math-3kInspired by @OpenR1
The questions for this dataset were all sourced from the first 3.3k prompts in open-r1/OpenR1-Math-220k
Dataset Stats (provided by OpenRouter):
Cost: $ 21.1 (USD)
Tokens (input + output): 52.3 M
glm-4.7-2000xThis is a reasoning dataset created using GLM 4.7 with reasoning effort set to high (not sure if that flag does anything for this model though).
The dataset is meant for creating distilled versions of GLM 4.7 by fine-tuning already existing open-source LLMs.
This dataset includes an addition to our recently enhanced set of prompts to cover creative writing multilingual reasoning and various graduate level questions across a wide variety of domains.
Stats
Costs: $ 30.72 (USD)… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/glm-4.7-2000x.
