CoolFace
20 results

capability

Hkang /moda-general-capability-rollouts MODA General Capability Retention Rollouts This dataset contains the raw model generations and evaluation results for the MODA general-capability retention experiments. It covers 16 models, seven benchmarks, 260,592 prompt records, and 2,605,920 stored generations. The evaluation code is pinned to source commit 12ea99b2a57a354f2b7d6792f62a3d9313192fa7. Evaluation protocol Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ, HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.tabulartext-generation1M<n<10M0 likes582 downloads1mo agoHugging Facesangamdas /YOUR-PHONE-NUMBER-IS-A-BEARER-TOKEN-CVID-CAPABILITY-BASED-REACHABILITY Your Phone Number Is a Bearer Token CVID — Capability-Validated Inbound Descriptors. Reachability as consumable, revocable, cryptographically bound authority instead of persistent exposure. Possession of an identifier is not permission to reach. Abstract A basic architectural weakness remains embedded in modern telecommunications: knowing how to reach someone is often practically equivalent to having permission to attempt contact. A phone number, email address… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/YOUR-PHONE-NUMBER-IS-A-BEARER-TOKEN-CVID-CAPABILITY-BASED-REACHABILITY.document10K<n<100K1 likes476 downloads21d agoHugging FacerealzL /benchability-fig4-capability-guided BenchAbility Figure 4 -- capability_guided One of two training mixtures drawn from the same frozen 884,143-row candidate pool, with the same budget (60,000 intervention + 15,000 shared replay) and the same hyperparameters. The two differ only in how the samples are chosen, which is the whole experiment. arm capability_guided selection by capability, gap-weighted from the Figure 3 scores intervention rows 59,999 replay rows 15,000 shards 38 pool 884,143… See the full description on the dataset page: https://huggingface.co/datasets/realzL/benchability-fig4-capability-guided.textvisual-question-answering10K<n<100K0 likes309 downloads1mo agoHugging Facet2ance /mat-02-9b-capability-ceiling 02 9B capability ceiling Is Qwen3.5-9B (19.3 GB of bf16 weights) a viable fast-iteration platform for the project's tree-search GRPO training, in place of the much larger Qwen3.6-27B (54 GB) student, without losing so much task capability that a cheaper training step stops being a useful learning iteration? The report's answer, quoting its abstract: "on these four tasks the 9B is not a viable fast platform" -- per node a 9B training step was 2.7 to 3.0x cheaper than the matching… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/mat-02-9b-capability-ceiling.0 likes218 downloads16d agoHugging Facedougalldeepmind /2026-07-31-qwen36-27b-capability-eval-arena-hard Qwen3.6-27B capability regression eval — Arena-Hard-v2.0 SxS (2026-07-31) Side-by-side capability eval for the synthetic-constitution-document SFT dose-response ladder on Qwen3.6-27B. Candidate answers generated with vLLM 0.26 (temperature 0, thinking on, identical decoding across arms); pairwise judging by google/gemini-3-flash-preview via OpenRouter against the arm_b (90/10) baseline, using the patched arena-hard-auto harness (baseline override, usage capture).… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-27b-capability-eval-arena-hard.0 likes208 downloads21d agoHugging Facedougalldeepmind /2026-07-31-qwen36-27b-mmlu-capability-eval 2026-07-31 — MMLU capability eval: Qwen3.6-27B constitution-SFT arm ladder experiment: Absolute-benchmark (MMLU) capability check that mixing synthetic constitution / difficult-advice documents into a Tulu SFT mixture does not cost Qwen3.6-27B general knowledge — the guardrail under the alignment result, run across the full mixture-ratio arm ladder against the untuned base model. date_generated: 2026-07-31 (think/, primary) and 2026-07-30 (nothink/, companion run) constitution:… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-27b-mmlu-capability-eval.0 likes189 downloads21d agoHugging Face