CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TeichAI /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K106 likes5.4k downloads4mo agoHugging Face02TeichAI /Ox-Alpha-Pi-TracesThis dataset was generated using teich by TeichAI Ox-Alpha Pi Agent Coding Traces This directory contains raw agent trace files generated by teich. JSONL files: 2247 Model metadata: stealth/ox-alpha Domains and prompt distribution Topic Traces Games & simulation (headless) 196 Frontend & Node-testable web 159 Health & medicine informatics 139 ML & scientific computing (CPU) 123 Data analysis & reporting 122 Computational biology & chemistry… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Ox-Alpha-Pi-Traces.text-generation8 likes3.1k downloads28d agoHugging Face03TeichAI /Ox-Alpha-10k Ox Alpha - 10k 10,005 single-turn prompts for text-response teacher generation Each row carries id, category, subcategory All data was gathered using stealth/ox-alpha via OpenRouter (reasoning effort high) Topic distribution Category Rows Share Coding (incl. Go/Rust, C++/Java/C#, shell/CLI) 944 9.5% Knowledge QA 891 9.0% Logical reasoning & decisions 734 7.4% Web development 720 7.2% Game development 720 7.2% Three.js / browser 3D 620 6.2%… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Ox-Alpha-10k.texttext-generation1K<n<10K34 likes848 downloads28d agoHugging Face04TeichAI /Fable-5-Cursor-TracesFable 5 Cursor Traces 244 Fable 5 Cursor agent sessions for training & research. This dataset has 244 Cursor sessions with Fable 5 at High/xHigh/Max effort levels for distillation. [!IMPORTANT] This dataset is compatible with Teich! Use it directly in your Teich training pipeline. [!WARNING] The longest rows exceed one million characters of content. Apply prepare_data() with your intended tokenizer and an explicit context/oversize policy before training.… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Fable-5-Cursor-Traces.textn<1K17 likes660 downloads25d agoHugging Face05TeichAI /claude-4.5-opus-high-reasoning-250xThis is a reasoning dataset created using Claude Opus 4.5 with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of Claude Opus 4.5 by fine-tuning already existing open-source LLMs. Stats Costs: $ 52.3 (USD) Total tokens (input + output): 2.13 M textn<1K404 likes479 downloads10mo agoHugging Face06armand0e /teich-test-v1 hy3-preview coding agent traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by tencent/hy3-preview:free. Training-ready tools Use this tools payload when rendering converted examples through your training chat template. The same structure is emitted on each converted example as the tools field. [ { "type": "function", "function": { "name": "bash", "description": "Execute bash… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/teich-test-v1.tabularn<1K0 likes464 downloads5mo agoHugging Face07TeichAI /lordx64-claude-opus-4.7-max-cleaned reasoning-distill-claude-opus-4-7-max-cleaned Cleaned version of lordx64/reasoning-distill-claude-opus-4-7-max. See the original dataset for full provenance, collection methodology, and terms of use. Cleaning steps Step Filter Reason Rows removed 1 Simulated thinking (...) Rows with ... in thinking/response indicate the model learned to simulate reasoning (e.g., "Now I'm laying out the puzzle grids...") rather than actually performing it. This causes failures… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/lordx64-claude-opus-4.7-max-cleaned.text1K<n<10K24 likes378 downloads5mo agoHugging Face08TeichAI /Claude-Sonnet-4.6-Reasoning-1100x Claude Sonnet 4.6 - High Reasoning 1096 conversations, all single-turn user → assistant pairs created using Claude Sonnet 4.6 with reasoning effort set to high. This is a pure reasoning/critical-thinking distillation dataset. Heavily weighted toward analytical thinking across economics, ethics, public policy, psychology, and epistemology. Dataset Format { "messages": [ {"role": "user", "content": "..."}, {"role": "assistant", "thinking": "Reasoning..."… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Claude-Sonnet-4.6-Reasoning-1100x.text1K<n<10K47 likes327 downloads6mo agoHugging Face09TeichAI /Claude-Opus-4.6-Reasoning-887x Claude Opus 4.6 - High Reasoning This is a reasoning dataset generated using Claude Opus 4.6 with high reasoning effort It contains distilled reasoning traces from Bullshit Bench for bullshit detection, legal and life decisions data for generalization, traces for improving the models understanding of vague and lazy prompts and more. Formatting guide { "messages": [ {"role": "user", "content": "..."}, {"role": "assistant", "thinking": "...", "content": "Final… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Claude-Opus-4.6-Reasoning-887x.textn<1K91 likes216 downloads6mo agoHugging Face10TeichAI /DeepSeek-v4-Flash-ChatThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Teich Test This directory contains newline-delimited JSON training examples generated by teich. All assistant responses were generated by deepseek/deepseek-v4-flash. Rows: 6313 Format Each file is newline-delimited JSON where every line is already a training example. Chat-only datasets include messages… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Flash-Chat.text1K<n<10K9 likes204 downloads4mo agoHugging Face11TeichAI /gpt-5.2-high-reasoning-250x Generated using DataGen by TeichAI This is a reasoning dataset created using GPT 5.2 with a reasoning depth set to high. The dataset is meant for creating distilled versions of GPT 5.2 by fine-tuning already existing open-source LLMs. Stats Costs: $ 10.58 (USD) Total tokens (input + output): N\A textn<1K32 likes188 downloads9mo agoHugging Face12TeichAI /gemini-3-flash-preview Gemini 3 Flash Preview This is a reasoning dataset created using Gemini 3 Flash Preview with a reasoning depth set to high. The dataset is meant for creating distilled versions of Gemini 3 Flash Preview by fine-tuning already existing open-source LLMs. This dataset contains a collection of prompts categorized by themes such as benchmarks, psychology, web development, embedded systems, creative writing, design, finance, legal, marketing, and science. No real benchmark questions were… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/gemini-3-flash-preview.text10K<n<100K16 likes158 downloads9mo agoHugging Face13TeichAI /gpt-5.1-high-reasoning-1000xThis is a reasoning dataset created using GPT 5.1 (high reasoning) from OpenAI. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of GPT 5.1 (high reasoning) thinking by fine-tuning already existing open-source LLMs. Stats Costs: $ 36.4 (USD) Tokens: 3.94 M (total) text1K<n<10K35 likes153 downloads10mo agoHugging Face14TeichAI /convo-v1A conversation dataset generated using MiniMax M2.1 to help teach models how to talk to users. textn<1K10 likes140 downloads8mo agoHugging Face15TeichAI /gpt-5-codex-1000xThis is a reasoning dataset created using GPT 5 Codex from OpenAI. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of GPT 5 Codex by fine-tuning already existing open-source LLMs. textn<1K8 likes125 downloads10mo agoHugging Face16TeichAI /claude-haiku-4.5-high-reasoning-1700xThis is a reasoning dataset created using Claude Haiku 4.5 with reasoning effort set to high. The dataset is meant for creating distilled versions of Claude Haiku 4.5 by fine-tuning already existing open-source LLMs. This dataset includes an addition to our recently enhanced set of prompts to cover creative writing and multilingual creative writing. Stats Costs: $ 33.52 (USD) Total tokens (input + output): 6.79 M text1K<n<10K7 likes115 downloads9mo agoHugging Face17agentlans /TeichAI-thinking-reasoning-x TeichAI Thinking & Reasoning Datasets A collection of prompts answered by large language models (LLMs) such as Google Gemini and OpenAI ChatGPT, with long-form reasoning enabled. These datasets were originally created by TeichAI for distillation and reasoning-focused training workflows. Schema Each row in the dataset has the following fields: question_hash: Truncated, base64-encoded MD5 hash of the question, useful for filtering and deduplication. question: The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/TeichAI-thinking-reasoning-x.texttext-generation100K<n<1M1 likes112 downloads5mo agoHugging Face18TeichAI /gemini-3-pro-preview-high-reasoning-1000xThis is a reasoning dataset created using Gemini 3 Pro Preview with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of Gemini 3 Pro Preview by fine-tuning already existing open-source LLMs on the summarized reasoning traces provided from their API. This dataset includes 250x from TeichAI/gemini-3-pro-preview-high-reasoning-250x Stats Costs: $ 32.7 (USD) Tokens: 2.73 M… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/gemini-3-pro-preview-high-reasoning-1000x.text1K<n<10K79 likes93 downloads10mo agoHugging Face19TeichAI /Gemini-3-Flash-Preview-VIBE Gemini 3 Flash Preview VIBE This dataset is our first attempt at an agentic coding SFT dataset. All of the prompts for this dataset were sourced from MiniMaxAI/VIBE. Each prompt was given to Gemini 3 Flash Preview with the follow tools and system prompt: read_file - Read file contents from workspace write_file - Write content to a file edit_file - Replace text in a file list_directory - List files and directories search_code - Search for patterns in files run_command - Execute… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Gemini-3-Flash-Preview-VIBE.textn<1K5 likes93 downloads8mo agoHugging Face20TeichAI /Claude-Opus-Dataclaw-Unredacted Claude Opus Dataclaw Unredacted How this dataset was built Collected the local Petromallet raw export plus selected public Dataclaw uploads. Filtered to the supported Opus-family source rows. Deduplicated by session_id and first user message. Converted raw assistant tool_uses directly into structured OpenAI-style tool_calls. Derived per-row tool definitions from canonical schemas and observed tool usage. Preserved assistant reasoning in <think>...</think> blocks.… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Claude-Opus-Dataclaw-Unredacted.texttext-generationn<1K22 likes91 downloads6mo agoHugging Face21mondk /TeichAI-ClaudeOpus4.5-High-CLEANEDNote: Base datasets: TeichAI/claude-4.5-opus-high-reasoning-250x Cleaned: low-quality data removed and content condensed (8MB -> 5MB) It cost me $2 (USD) [just kidding] textn<1K5 likes91 downloads1mo agoHugging Face22TeichAI /MiniMax-M2.1-Code-SFT MiniMax M2.1 Code SFT 200 of the prompts for this dataset were sourced from MiniMaxAI/VIBE. The rest were generated. Each prompt was given to MiniMax M2.1 with the follow tools and system prompt: read_file - Read file contents from workspace write_file - Write content to a file edit_file - Replace text in a file list_directory - List files and directories search_code - Search for patterns in files run_command - Execute shell commands (with timeout) web_search - Web search (powered… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/MiniMax-M2.1-Code-SFT.text1K<n<10K18 likes87 downloads8mo agoHugging Face23TeichAI /grok-code-fast-1-1000xThis is a reasoning dataset created using Grok Code Fast 1 from xAI. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of Grok Code Fast 1 by fine-tuning already existing open-source LLMs. text1K<n<10K6 likes86 downloads1y agoHugging Face24TeichAI /gemini-3-pro-preview-high-reasoning-250xThis is a reasoning dataset created using Gemini 3 Pro Preview with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of Gemini 3 Pro Preview by fine-tuning already existing open-source LLMs on the summarized reasoning traces provided from their API. Stats Costs: $ 5.78 Total tokens (input + output): 484 K textn<1K42 likes84 downloads8mo agoHugging Face25TeichAI /claude-sonnet-4.5-high-reasoning-250xThis is a reasoning dataset created using Claude Sonnet 4.5 with a high reasoning effort. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of Claude Sonnet 4.5 by fine-tuning already existing open-source LLMs. The default system prompt from OpenrouterAI was used You are Claude Sonnet 4.5, a large language model from anthropic. Formatting Rules: - Use Markdown for lists, tables, and styling. - Use ```code fences```… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/claude-sonnet-4.5-high-reasoning-250x.textn<1K39 likes83 downloads11mo agoHugging Face26TeichAI /Pony-Alpha-15k Pony Alpha 15k This is a reasoning dataset generated using the stealth model Pony Alpha, which ended up being GLM-5. As the largest dataset we have made yet. The prompts from this dataset were almost all generated by GPT 5.1 and Gemini 3 (flash and pro). The categories covered include academia, multi-lingual creative writing, finance, health, law, marketing/SEO, programming, philosophy, web dev, python scripting, and science. Stats: Cost: $ 0 (USD) Tokens (input + output): 43.3 M text10K<n<100K66 likes83 downloads7mo agoHugging Face27TeichAI /polaris-alpha-1000xThis is a non-reasoning dataset created using Polaris Alpha (likely an alpha version of a new ChatGPT) from OpenAI. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of Polaris Alpha by fine-tuning already existing open-source LLMs. text1K<n<10K11 likes80 downloads11mo agoHugging Face28TeichAI /gpt-5.1-codex-max-1000x GPT 5.1 Codex Max - 1,000x This is a reasoning dataset created using GPT 5.1 Codex Max with a reasoning depth set to high. The dataset is meant for creating distilled versions of GPT 5.1 Codex Max by fine-tuning already existing open-source LLMs. This datasets prompts are adapted versions of our standard set of prompts, designed to provoke complex reasoning. Stats Costs: $ 23.51 (USD) Total tokens (input + output): 2.39 M Generated using DataGen by TeichAI text1K<n<10K27 likes70 downloads9mo agoHugging Face29TeichAI /mistral-small-creative-500x Mistral Small Creative - 500x This is a non-reasoning dataset created using Mistral Small Creative. The dataset is meant for creating distilled versions of Mistral Small Creative by fine-tuning already existing open-source LLMs. This dataset only covers generating short fictional stories. Meant for testing purposes -> to evaluate if creating a larger dataset would be worth it Stats Costs: $ 0.20 (USD) Total tokens (input + output): 704K textn<1K8 likes70 downloads8mo agoHugging Face30TeichAI /kimi-k2-thinking-1000xThis is a reasoning dataset created using Kimi k2 thinking from MoonshotAI. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of Kimi k2 thinking by fine-tuning already existing open-source LLMs. textn<1K14 likes65 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.