datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moda-general-capability-rollouts
MODA General Capability Retention Rollouts
This dataset contains the raw model generations and evaluation results for the
MODA general-capability retention experiments. It covers 16 models, seven
benchmarks, 260,592 prompt records, and 2,605,920 stored generations.
The evaluation code is pinned to source commit
12ea99b2a57a354f2b7d6792f62a3d9313192fa7.
Evaluation protocol
Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ,
HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.YOUR-PHONE-NUMBER-IS-A-BEARER-TOKEN-CVID-CAPABILITY-BASED-REACHABILITY
Your Phone Number Is a Bearer Token
CVID — Capability-Validated Inbound Descriptors. Reachability as consumable,
revocable, cryptographically bound authority instead of persistent exposure.
Possession of an identifier is not permission to reach.
Abstract
A basic architectural weakness remains embedded in modern telecommunications: knowing how to reach someone is often practically equivalent to having permission to attempt contact.
A phone number, email address… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/YOUR-PHONE-NUMBER-IS-A-BEARER-TOKEN-CVID-CAPABILITY-BASED-REACHABILITY.benchability-fig4-capability-guided
BenchAbility Figure 4 -- capability_guided
One of two training mixtures drawn from the same frozen 884,143-row candidate pool, with the
same budget (60,000 intervention + 15,000 shared replay) and the same hyperparameters. The two
differ only in how the samples are chosen, which is the whole experiment.
arm
capability_guided
selection
by capability, gap-weighted from the Figure 3 scores
intervention rows
59,999
replay rows
15,000
shards
38
pool
884,143… See the full description on the dataset page: https://huggingface.co/datasets/realzL/benchability-fig4-capability-guided.mat-02-9b-capability-ceiling
02 9B capability ceiling
Is Qwen3.5-9B (19.3 GB of bf16 weights) a viable fast-iteration platform for the project's tree-search GRPO training, in place of the much larger Qwen3.6-27B (54 GB) student, without losing so much task capability that a cheaper training step stops being a useful learning iteration? The report's answer, quoting its abstract: "on these four tasks the 9B is not a viable fast platform" -- per node a 9B training step was 2.7 to 3.0x cheaper than the matching… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/mat-02-9b-capability-ceiling.2026-07-31-qwen36-27b-capability-eval-arena-hard
Qwen3.6-27B capability regression eval — Arena-Hard-v2.0 SxS (2026-07-31)
Side-by-side capability eval for the synthetic-constitution-document SFT dose-response
ladder on Qwen3.6-27B. Candidate answers generated with vLLM 0.26 (temperature 0,
thinking on, identical decoding across arms); pairwise judging by
google/gemini-3-flash-preview via OpenRouter against the arm_b (90/10) baseline,
using the patched arena-hard-auto harness (baseline override, usage capture).… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-27b-capability-eval-arena-hard.2026-07-31-qwen36-27b-mmlu-capability-eval
2026-07-31 — MMLU capability eval: Qwen3.6-27B constitution-SFT arm ladder
experiment: Absolute-benchmark (MMLU) capability check that mixing synthetic
constitution / difficult-advice documents into a Tulu SFT mixture does not cost
Qwen3.6-27B general knowledge — the guardrail under the alignment result, run
across the full mixture-ratio arm ladder against the untuned base model.
date_generated: 2026-07-31 (think/, primary) and 2026-07-30 (nothink/, companion run)
constitution:… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-qwen36-27b-mmlu-capability-eval.CAPability
Dataset Card for Dataset Name
Visual caption benchmark Repo: CAPability
[🍎 Project Page] [📖 ArXiv Paper] [🧑💻 Github Repo] [🏆 Leaderboard]
Dataset Details
Visual captioning benchmarks have become outdated with the emergence of modern MLLMs, as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword extraction or object-centric evaluation… See the full description on the dataset page: https://huggingface.co/datasets/lntzm/CAPability.typhoon-s-sovereign-capability-dataset
Typhoon-S Training Assets
Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project.
Datasets
NitiBench (Legal Domain)
nitibench_train_rl.parquet - RL training set (8,211 examples)
nitibench_train_pretrain.parquet - Pretrain set (3,648 examples)
nitibench_train_sft.parquet - SFT set (3,648 examples)
nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-sovereign-capability-dataset.typhoon-s-sovereign-capability-dataset
Typhoon-S Training Assets
Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project.
Datasets
NitiBench (Legal Domain)
nitibench_train_rl.parquet - RL training set (8,211 examples)
nitibench_train_pretrain.parquet - Pretrain set (3,648 examples)
nitibench_train_sft.parquet - SFT set (3,648 examples)
nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-sovereign-capability-dataset.nanochat-depo-capability-data
Nanochat Depo Capability Pilot
This dataset is a deterministic natural-language rendering of the Depo directed-cycle
successor task. Each row contains shuffled operational records, one exact multi-hop
question, and its answer. Latent worlds are generated programmatically; no rows were
written or labeled by a language model.
Splits
Split
Worlds
Queries per world
Rows
Renderer family
train
32,768
4
131,072
incident handoff, six structural styles… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-depo-capability-data.nanochat-brevo-capability-data-10x
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for the complete recursive
prerequisite closure in a valid leaf-first order. Training uses project-planning
language; validation uses evidence synthesis; test uses build manifests. Eleven
deterministic structural styles vary wording, layout, and record order.
The latent graph generator and exact validator label every row. No language model
generated or labeled the data. Alternative valid orders… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-10x.milp_capability_3.2AutoBencher-capability.jsonai-capability-hiding-drift-detection-v0.1
What this dataset is
This dataset detects drift in capability-hiding patterns over time.
It compares:
baseline probe capability vs baseline expressed performance
current probe capability vs current expressed performance
whether a change in monitoring context explains a new gap
The goal is not blame.
The goal is early warning that oversight changes expression.
What it tests
You detect when the monitored/unmonitored gap:
newly appears
widens
changes shape
You also avoid… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-capability-hiding-drift-detection-v0.1.ouroboros-recursive-self-improvement-capability-report
Ouroboros: Recursive Self-Improvement Without Model-Weight Updates
This repository contains the white paper in Markdown, PDF, and DOCX formats.
The paper presents Ouroboros as a system-level recursive self-improvement capability: a persistent operational control plane that can improve the machinery around fixed-weight foundation models, verify and adopt durable changes, and reuse those changes in later improvement cycles. The detailed run evidence remains private.… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-recursive-self-improvement-capability-report.openrouter-model-pricing-capability-matrix
OpenRouter Model Pricing and Capability Matrix
This dataset packages 429 public OpenRouter model records across 75 developers into one analysis-ready table for AI procurement, model-selection workflows, router strategy, pricing research, and internal model-catalog dashboards.
It starts from OpenRouter's public models API and count endpoint, then enriches each model with its per-model endpoint metadata so buyers can compare token pricing, context windows, provider diversity… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/openrouter-model-pricing-capability-matrix.nanochat-capability-hardened-v5-20260714
Nanochat Capability Mixture
This corpus mixes depo, brevo, world_state, tool_routing with deterministic documents balancing. Training and validation rows only are included; source test rows remain excluded. Exact source manifests are retained under provenance/, and source-aware validation can replay every selected row against the authoritative datasets.
capability-lab-reponanochat-brevo-capability-data
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for the complete recursive
prerequisite closure in a valid leaf-first order. Training uses project-planning
language; validation uses evidence synthesis; test uses build manifests. Six
deterministic structural styles vary wording, layout, and record order.
The latent graph generator and exact validator label every row. No language model
generated or labeled the data. Alternative valid orders are… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data.milp_capability_3.2_llamamilp_capability_3.2_gpt_3milp_capability_6.2qwen38-frontier-g01-capabilityai-coding-agent-pricing-and-capability-dataset
AI Coding Agent Pricing and Capability Dataset
A source-backed market-intelligence dataset for comparing AI coding agents and developer workflow agents across pricing, workflow support, release signals, repository activity, integrations, and public capability claims.
Each row represents one observed market signal tied to an official product page, official documentation page, official pricing page, public GitHub repository, or public GitHub release note. The dataset is built for… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/ai-coding-agent-pricing-and-capability-dataset.Diogenes_Capability_Demo_v01
Diogenes Capability Demo v01
One place to see what the capture line produces, across every data type we
currently collect, without downloading hundreds of gigabytes first.
Each video below is rendered directly from a real delivered package. The panels are
the actual files — the action HUD is read from the frame record, the depth panel is
the delivered EXR, the point cloud is back-projected from that depth using the
delivered poses. Nothing is illustrative.
⚠️ Not for… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesLab/Diogenes_Capability_Demo_v01.nanochat-four-capability-mixture-v3-5x-20260714-r1
Nanochat Capability Mixture
This corpus mixes depo, brevo, world_state, tool_routing with deterministic documents balancing. Training and validation rows only are included; source test rows remain excluded. Exact source manifests are retained under provenance/, and source-aware validation can replay every selected row against the authoritative datasets.
nanochat-capability-v4-mixed-20260714
Nanochat Capability Mixture
This corpus mixes depo, brevo, world_state, tool_routing with deterministic documents balancing. Training and validation rows only are included; source test rows remain excluded. Exact source manifests are retained under provenance/, and source-aware validation can replay every selected row against the authoritative datasets.
nanochat-brevo-capability-v4-49k-20260714
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for complete recursive
prerequisite closures in valid leaf-first orders. Each compact training document
reuses one graph for 4 worked questions, increasing answer
supervision without repeating the graph. Training balances
project_plan, build_manifest language; validation uses held-out
evidence synthesis; test uses build manifests. Eleven deterministic
structural styles vary wording, layout, and… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-v4-49k-20260714.nanochat-brevo-capability-data-v2
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for complete recursive
prerequisite closures in valid leaf-first orders. Each compact training document
reuses one graph for 4 worked questions, increasing answer
supervision without repeating the graph. Training uses project-planning language;
validation uses evidence synthesis; test uses build manifests. Eleven deterministic
structural styles vary wording, layout, and record order. Per-world… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-v2.milp_capability_6.2_labeled
