CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TeichAI /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K108 likes6.1k downloads4mo agoHugging Face02TeichAI /Ox-Alpha-10k Ox Alpha - 10k 10,005 single-turn prompts for text-response teacher generation Each row carries id, category, subcategory All data was gathered using stealth/ox-alpha via OpenRouter (reasoning effort high) Topic distribution Category Rows Share Coding (incl. Go/Rust, C++/Java/C#, shell/CLI) 944 9.5% Knowledge QA 891 9.0% Logical reasoning & decisions 734 7.4% Web development 720 7.2% Game development 720 7.2% Three.js / browser 3D 620 6.2%… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Ox-Alpha-10k.texttext-generation1K<n<10K34 likes858 downloads1mo agoHugging Face03TeichAI /Fable-5-Cursor-TracesFable 5 Cursor Traces 244 Fable 5 Cursor agent sessions for training & research. This dataset has 244 Cursor sessions with Fable 5 at High/xHigh/Max effort levels for distillation. [!IMPORTANT] This dataset is compatible with Teich! Use it directly in your Teich training pipeline. [!WARNING] The longest rows exceed one million characters of content. Apply prepare_data() with your intended tokenizer and an explicit context/oversize policy before training.… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Fable-5-Cursor-Traces.textn<1K17 likes700 downloads27d agoHugging Face04TeichAI /claude-4.5-opus-high-reasoning-250xThis is a reasoning dataset created using Claude Opus 4.5 with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of Claude Opus 4.5 by fine-tuning already existing open-source LLMs. Stats Costs: $ 52.3 (USD) Total tokens (input + output): 2.13 M textn<1K404 likes481 downloads10mo agoHugging Face05armand0e /teich-test-v1 hy3-preview coding agent traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by tencent/hy3-preview:free. Training-ready tools Use this tools payload when rendering converted examples through your training chat template. The same structure is emitted on each converted example as the tools field. [ { "type": "function", "function": { "name": "bash", "description": "Execute bash… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/teich-test-v1.tabularn<1K0 likes464 downloads5mo agoHugging Face06TeichAI /lordx64-claude-opus-4.7-max-cleaned reasoning-distill-claude-opus-4-7-max-cleaned Cleaned version of lordx64/reasoning-distill-claude-opus-4-7-max. See the original dataset for full provenance, collection methodology, and terms of use. Cleaning steps Step Filter Reason Rows removed 1 Simulated thinking (...) Rows with ... in thinking/response indicate the model learned to simulate reasoning (e.g., "Now I'm laying out the puzzle grids...") rather than actually performing it. This causes failures… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/lordx64-claude-opus-4.7-max-cleaned.text1K<n<10K24 likes397 downloads5mo agoHugging Face07TeichAI /DeepSeek-v4-Flash-ChatThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Teich Test This directory contains newline-delimited JSON training examples generated by teich. All assistant responses were generated by deepseek/deepseek-v4-flash. Rows: 6313 Format Each file is newline-delimited JSON where every line is already a training example. Chat-only datasets include messages… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Flash-Chat.text1K<n<10K9 likes194 downloads5mo agoHugging Face08TeichAI /gpt-5.2-high-reasoning-250x Generated using DataGen by TeichAI This is a reasoning dataset created using GPT 5.2 with a reasoning depth set to high. The dataset is meant for creating distilled versions of GPT 5.2 by fine-tuning already existing open-source LLMs. Stats Costs: $ 10.58 (USD) Total tokens (input + output): N\A textn<1K32 likes188 downloads10mo agoHugging Face09TeichAI /gemini-3-flash-preview Gemini 3 Flash Preview This is a reasoning dataset created using Gemini 3 Flash Preview with a reasoning depth set to high. The dataset is meant for creating distilled versions of Gemini 3 Flash Preview by fine-tuning already existing open-source LLMs. This dataset contains a collection of prompts categorized by themes such as benchmarks, psychology, web development, embedded systems, creative writing, design, finance, legal, marketing, and science. No real benchmark questions were… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/gemini-3-flash-preview.text10K<n<100K16 likes174 downloads9mo agoHugging Face10TeichAI /gpt-5.1-high-reasoning-1000xThis is a reasoning dataset created using GPT 5.1 (high reasoning) from OpenAI. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of GPT 5.1 (high reasoning) thinking by fine-tuning already existing open-source LLMs. Stats Costs: $ 36.4 (USD) Tokens: 3.94 M (total) text1K<n<10K35 likes155 downloads10mo agoHugging Face11TeichAI /convo-v1A conversation dataset generated using MiniMax M2.1 to help teach models how to talk to users. textn<1K10 likes137 downloads9mo agoHugging Face12agentlans /TeichAI-thinking-reasoning-x TeichAI Thinking & Reasoning Datasets A collection of prompts answered by large language models (LLMs) such as Google Gemini and OpenAI ChatGPT, with long-form reasoning enabled. These datasets were originally created by TeichAI for distillation and reasoning-focused training workflows. Schema Each row in the dataset has the following fields: question_hash: Truncated, base64-encoded MD5 hash of the question, useful for filtering and deduplication. question: The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/TeichAI-thinking-reasoning-x.texttext-generation100K<n<1M1 likes127 downloads5mo agoHugging Face13TeichAI /claude-haiku-4.5-high-reasoning-1700xThis is a reasoning dataset created using Claude Haiku 4.5 with reasoning effort set to high. The dataset is meant for creating distilled versions of Claude Haiku 4.5 by fine-tuning already existing open-source LLMs. This dataset includes an addition to our recently enhanced set of prompts to cover creative writing and multilingual creative writing. Stats Costs: $ 33.52 (USD) Total tokens (input + output): 6.79 M text1K<n<10K7 likes125 downloads9mo agoHugging Face14TeichAI /gpt-5-codex-1000xThis is a reasoning dataset created using GPT 5 Codex from OpenAI. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of GPT 5 Codex by fine-tuning already existing open-source LLMs. textn<1K8 likes120 downloads11mo agoHugging Face15TeichAI /Pony-Alpha-15k Pony Alpha 15k This is a reasoning dataset generated using the stealth model Pony Alpha, which ended up being GLM-5. As the largest dataset we have made yet. The prompts from this dataset were almost all generated by GPT 5.1 and Gemini 3 (flash and pro). The categories covered include academia, multi-lingual creative writing, finance, health, law, marketing/SEO, programming, philosophy, web dev, python scripting, and science. Stats: Cost: $ 0 (USD) Tokens (input + output): 43.3 M text10K<n<100K66 likes97 downloads7mo agoHugging Face16TeichAI /Gemini-3-Flash-Preview-VIBE Gemini 3 Flash Preview VIBE This dataset is our first attempt at an agentic coding SFT dataset. All of the prompts for this dataset were sourced from MiniMaxAI/VIBE. Each prompt was given to Gemini 3 Flash Preview with the follow tools and system prompt: read_file - Read file contents from workspace write_file - Write content to a file edit_file - Replace text in a file list_directory - List files and directories search_code - Search for patterns in files run_command - Execute… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Gemini-3-Flash-Preview-VIBE.textn<1K5 likes93 downloads8mo agoHugging Face17TeichAI /Claude-Opus-Dataclaw-Unredacted Claude Opus Dataclaw Unredacted How this dataset was built Collected the local Petromallet raw export plus selected public Dataclaw uploads. Filtered to the supported Opus-family source rows. Deduplicated by session_id and first user message. Converted raw assistant tool_uses directly into structured OpenAI-style tool_calls. Derived per-row tool definitions from canonical schemas and observed tool usage. Preserved assistant reasoning in <think>...</think> blocks.… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Claude-Opus-Dataclaw-Unredacted.texttext-generationn<1K22 likes91 downloads6mo agoHugging Face18TeichAI /gemini-3-pro-preview-high-reasoning-1000xThis is a reasoning dataset created using Gemini 3 Pro Preview with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of Gemini 3 Pro Preview by fine-tuning already existing open-source LLMs on the summarized reasoning traces provided from their API. This dataset includes 250x from TeichAI/gemini-3-pro-preview-high-reasoning-250x Stats Costs: $ 32.7 (USD) Tokens: 2.73 M… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/gemini-3-pro-preview-high-reasoning-1000x.text1K<n<10K79 likes90 downloads10mo agoHugging Face19TeichAI /grok-code-fast-1-1000xThis is a reasoning dataset created using Grok Code Fast 1 from xAI. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of Grok Code Fast 1 by fine-tuning already existing open-source LLMs. text1K<n<10K6 likes87 downloads1y agoHugging Face20TeichAI /claude-sonnet-4.5-high-reasoning-250xThis is a reasoning dataset created using Claude Sonnet 4.5 with a high reasoning effort. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of Claude Sonnet 4.5 by fine-tuning already existing open-source LLMs. The default system prompt from OpenrouterAI was used You are Claude Sonnet 4.5, a large language model from anthropic. Formatting Rules: - Use Markdown for lists, tables, and styling. - Use ```code fences```… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/claude-sonnet-4.5-high-reasoning-250x.textn<1K39 likes84 downloads11mo agoHugging Face21TeichAI /gemini-3-pro-preview-high-reasoning-250xThis is a reasoning dataset created using Gemini 3 Pro Preview with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of Gemini 3 Pro Preview by fine-tuning already existing open-source LLMs on the summarized reasoning traces provided from their API. Stats Costs: $ 5.78 Total tokens (input + output): 484 K textn<1K42 likes84 downloads8mo agoHugging Face22TeichAI /MiniMax-M2.1-Code-SFT MiniMax M2.1 Code SFT 200 of the prompts for this dataset were sourced from MiniMaxAI/VIBE. The rest were generated. Each prompt was given to MiniMax M2.1 with the follow tools and system prompt: read_file - Read file contents from workspace write_file - Write content to a file edit_file - Replace text in a file list_directory - List files and directories search_code - Search for patterns in files run_command - Execute shell commands (with timeout) web_search - Web search (powered… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/MiniMax-M2.1-Code-SFT.text1K<n<10K18 likes84 downloads8mo agoHugging Face23TeichAI /polaris-alpha-1000xThis is a non-reasoning dataset created using Polaris Alpha (likely an alpha version of a new ChatGPT) from OpenAI. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of Polaris Alpha by fine-tuning already existing open-source LLMs. text1K<n<10K11 likes79 downloads11mo agoHugging Face24mondk /TeichAI-ClaudeOpus4.5-High-CLEANEDNote: Base datasets: TeichAI/claude-4.5-opus-high-reasoning-250x Cleaned: low-quality data removed and content condensed (8MB -> 5MB) It cost me $2 (USD) [just kidding] textn<1K5 likes78 downloads2mo agoHugging Face25TeichAI /mistral-small-creative-500x Mistral Small Creative - 500x This is a non-reasoning dataset created using Mistral Small Creative. The dataset is meant for creating distilled versions of Mistral Small Creative by fine-tuning already existing open-source LLMs. This dataset only covers generating short fictional stories. Meant for testing purposes -> to evaluate if creating a larger dataset would be worth it Stats Costs: $ 0.20 (USD) Total tokens (input + output): 704K textn<1K8 likes74 downloads8mo agoHugging Face26TeichAI /kimi-k2-thinking-1000xThis is a reasoning dataset created using Kimi k2 thinking from MoonshotAI. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of Kimi k2 thinking by fine-tuning already existing open-source LLMs. textn<1K14 likes70 downloads11mo agoHugging Face27TeichAI /gpt-5.1-codex-max-1000x GPT 5.1 Codex Max - 1,000x This is a reasoning dataset created using GPT 5.1 Codex Max with a reasoning depth set to high. The dataset is meant for creating distilled versions of GPT 5.1 Codex Max by fine-tuning already existing open-source LLMs. This datasets prompts are adapted versions of our standard set of prompts, designed to provoke complex reasoning. Stats Costs: $ 23.51 (USD) Total tokens (input + output): 2.39 M Generated using DataGen by TeichAI text1K<n<10K27 likes68 downloads9mo agoHugging Face28TeichAI /Hunter-Alpha-16k Hunter Alpha 16k This is a reasoning dataset generated using the stealth model Hunter Alpha, which was revealed to be xiaomi/mimo-v2-pro. As the largest dataset we have made yet. The prompts from this dataset were almost all generated by Opus 4.5/4.6, GPT 5.1 and Gemini 3 (flash and pro). The categories covered include academia, multi-lingual creative writing, finance, health, law, marketing/SEO, programming, philosophy, web dev, python scripting, and science. Stats: Cost: $ 0… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Hunter-Alpha-16k.text10K<n<100K8 likes66 downloads6mo agoHugging Face29TeichAI /deepseek-v3.2-speciale-openr1-math-3kInspired by @OpenR1 The questions for this dataset were all sourced from the first 3.3k prompts in open-r1/OpenR1-Math-220k Dataset Stats (provided by OpenRouter): Cost: $ 21.1 (USD) Tokens (input + output): 52.3 M text1K<n<10K9 likes64 downloads10mo agoHugging Face30TeichAI /glm-4.7-2000xThis is a reasoning dataset created using GLM 4.7 with reasoning effort set to high (not sure if that flag does anything for this model though). The dataset is meant for creating distilled versions of GLM 4.7 by fine-tuning already existing open-source LLMs. This dataset includes an addition to our recently enhanced set of prompts to cover creative writing multilingual reasoning and various graduate level questions across a wide variety of domains. Stats Costs: $ 30.72 (USD)… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/glm-4.7-2000x.text1K<n<10K94 likes61 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.