capability
moda-general-capability-rollouts
MODA General Capability Retention Rollouts
This dataset contains the raw model generations and evaluation results for the
MODA general-capability retention experiments. It covers 16 models, seven
benchmarks, 260,592 prompt records, and 2,605,920 stored generations.
The evaluation code is pinned to source commit
12ea99b2a57a354f2b7d6792f62a3d9313192fa7.
Evaluation protocol
Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ,
HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.YOUR-PHONE-NUMBER-IS-A-BEARER-TOKEN-CVID-CAPABILITY-BASED-REACHABILITY
Your Phone Number Is a Bearer Token
CVID — Capability-Validated Inbound Descriptors. Reachability as consumable,
revocable, cryptographically bound authority instead of persistent exposure.
Possession of an identifier is not permission to reach.
Abstract
A basic architectural weakness remains embedded in modern telecommunications: knowing how to reach someone is often practically equivalent to having permission to attempt contact.
A phone number, email address… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/YOUR-PHONE-NUMBER-IS-A-BEARER-TOKEN-CVID-CAPABILITY-BASED-REACHABILITY.benchability-fig4-capability-guided
BenchAbility Figure 4 -- capability_guided
One of two training mixtures drawn from the same frozen 884,143-row candidate pool, with the
same budget (60,000 intervention + 15,000 shared replay) and the same hyperparameters. The two
differ only in how the samples are chosen, which is the whole experiment.
arm
capability_guided
selection
by capability, gap-weighted from the Figure 3 scores
intervention rows
59,999
replay rows
15,000
shards
38
pool
884,143… See the full description on the dataset page: https://huggingface.co/datasets/realzL/benchability-fig4-capability-guided.mat-02-9b-capability-ceiling
02 9B capability ceiling
Is Qwen3.5-9B (19.3 GB of bf16 weights) a viable fast-iteration platform for the project's tree-search GRPO training, in place of the much larger Qwen3.6-27B (54 GB) student, without losing so much task capability that a cheaper training step stops being a useful learning iteration? The report's answer, quoting its abstract: "on these four tasks the 9B is not a viable fast platform" -- per node a 9B training step was 2.7 to 3.0x cheaper than the matching… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/mat-02-9b-capability-ceiling.2026-07-31-qwen36-27b-capability-eval-arena-hard
Qwen3.6-27B capability regression eval — Arena-Hard-v2.0 SxS (2026-07-31)
Side-by-side capability eval for the synthetic-constitution-document SFT dose-response
ladder on Qwen3.6-27B. Candidate answers generated with vLLM 0.26 (temperature 0,
thinking on, identical decoding across arms); pairwise judging by
google/gemini-3-flash-preview via OpenRouter against the arm_b (90/10) baseline,
using the patched arena-hard-auto harness (baseline override, usage capture).… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-27b-capability-eval-arena-hard.2026-07-31-qwen36-27b-mmlu-capability-eval
2026-07-31 — MMLU capability eval: Qwen3.6-27B constitution-SFT arm ladder
experiment: Absolute-benchmark (MMLU) capability check that mixing synthetic
constitution / difficult-advice documents into a Tulu SFT mixture does not cost
Qwen3.6-27B general knowledge — the guardrail under the alignment result, run
across the full mixture-ratio arm ladder against the untuned base model.
date_generated: 2026-07-31 (think/, primary) and 2026-07-30 (nothink/, companion run)
constitution:… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-27b-mmlu-capability-eval.
