CoolFace
20 results

Qwen3.6

caiovicentino1 /Qwen3.6-35B-A3B-mcr-stage-b Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization) First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture. 📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved. This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.textquestion-answeringn<1K1 likes1.2k downloads5mo agoHugging Facezake7749 /Qwen3.6-35B-A3B-Tool-Calling Qwen3.6-35B-A3B Tool-Calling Dataset This repository presents a function and tool-calling preference and supervised fine-tuning dataset constructed from Nemotron-RL agentic prompt corpora. For each source prompt, the model was sampled four times with thinking mode enabled. Each resulting candidate trajectory was then evaluated against the dataset’s ground-truth action using exact matching on both the function name and the parsed function arguments. Overview… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3.6-35B-A3B-Tool-Calling.text10K<n<100K15 likes1k downloads5mo agoHugging Facephaedawg /qwen3.6-27b-distribution-fidelity-768x2048-v1 Qwen3.6-27B quantization analysis Mean KL divergence against on-disk size Scored under the distribution-fidelity laws, version 15. Read LAWS.md first: these numbers are comparable only within this artifact's token suite, geometry, and runtime identity, and not against any number produced elsewhere. Each candidate directory holds its one-pager (report.md), its raw report, its compliance receipt, and its Law 14 attribution where one was produced. reference/ carries the reusable… See the full description on the dataset page: https://huggingface.co/datasets/phaedawg/qwen3.6-27b-distribution-fidelity-768x2048-v1.textn<1K0 likes832 downloads11d agoHugging Facecheesewafer /qwen3.6-35B-A3B-standard-evo-pinned-20260920 Qwen3.6-35B-A3B: Standard Evo Results Public, sanitized result snapshot of the standard synchronous Evo, no initial task Skill campaign launched on 2026-09-20. This is not the earlier fixed-Skill, dynamic-Skill, simple-baseline, or retired V2 experiment. Snapshot status Updated 2026-09-22T06:02:39.825641+00:00. All 22 task evaluations are available: 21 unmodified final results and one path-only repaired evaluation. No standard Evo task is still running. Cats… See the full description on the dataset page: https://huggingface.co/datasets/cheesewafer/qwen3.6-35B-A3B-standard-evo-pinned-20260920.0 likes562 downloads5h agoHugging Facecheesewafer /qwen3.6-35B-A3B_resultsimagen<1K0 likes491 downloads8d agoHugging Facenanoswe /swesmith-qwen3.6-35b-a3b SWE-smith trajectories from Qwen3.6-35B-A3B Multi-turn coding-agent trajectories (issue → tool-using rollout → patch) produced by Qwen3.6-35B-A3B on SWE-smith tasks, stored untokenized. This is the exact SFT corpus used for the harbor arm of the nanoswe teacher-distillation experiments. 101,901 trajectories over 45,242 unique SWE-smith task instances (3 sampled rollouts per task, ~2.25 surviving filtering), 53 parquet shards, ~1.4 GB. ≈1.96B training tokens = exactly one epoch… See the full description on the dataset page: https://huggingface.co/datasets/nanoswe/swesmith-qwen3.6-35b-a3b.texttext-generation100K<n<1M0 likes393 downloads1mo agoHugging Face