thinking-traces
exp_tas_full_thinking_tracesexp_tas_interleaved_thinking_on_tracesQwen_Qwen3-4B-Thinking-2507_int3-g16-fp8_qwen3-traces-cot-concat_2048_64_1024_128_lr0.02staqc-sandboxes-traces-terminus-2_Qwen3-4B-Thinking-2507Qwen_Qwen3-4B-Thinking-2507_int4-g128_qwen3-traces-cot-concat_2048_8_1024_256_lr0.1Qwen_Qwen3-4B-Thinking-2507_PTQ_GPTQ_INT3-asym_qwen3-cot-tracesQwen_Qwen3-4B-Thinking-2507_PTQ_AUTOROUND_INT3-asym_qwen3-cot-tracesQwen_Qwen3-4B-Thinking-2507_PTQ_GPTQ_INT4-sym_qwen3-cot-traces
thinking-cap-tier-raw-traces
Thinking Cap Tier Raw Traces (TCS v4)
[!IMPORTANT]
Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture:
All 38,158 candidate reasoning traces across all 4 tiers (candidates_low.jsonl, candidates_mid.jsonl, candidates_high.jsonl, candidates_xhigh.jsonl) are 100% sanitized:
Zero batch-padding residues (<|pad|>): Completely purged across all records.
Strict Delimiter Integrity: Generation blocks cleanly separate thought deliberation tags… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-raw-traces.DCAgent2_terminal_bench_2_laion_exp_tas_full_thinking_traces_20260102_04565520260728-001051-qwen3-30b-a3b-thinking-2507-qwen-opencode-v2-a708-tracesDCAgent2_terminal_bench_2_laion_exp_tas_interleaved_thinking_on_traces_20260102_050258qwen35-a3b-thinking-traces
Qwen3.5-35B-A3B Thinking Traces — SAE Training Data
Per-sentence L17 residual activations from Qwen/Qwen3.5-35B-A3B generating CoT on MMLU-Pro.
Stats
Model: Qwen/Qwen3.5-35B-A3B
Layer: L17 residual (~42% depth of 40-layer hybrid MoE)
Prompts: 2000 from MMLU-Pro test
Sentences: 41285
d_model: 2048
Activation dtype: float16
Purpose
Replication of Venhoff et al. 2025 (arXiv:2510.07364) "Base Models Know How to Reason, Thinking Models Learn When" applied to… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/qwen35-a3b-thinking-traces.claude-4-5-sonnet-thinking-stackexchange-overflow-32ep-32k-traces
