datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moda-general-capability-rollouts
MODA General Capability Retention Rollouts
This dataset contains the raw model generations and evaluation results for the
MODA general-capability retention experiments. It covers 16 models, seven
benchmarks, 260,592 prompt records, and 2,605,920 stored generations.
The evaluation code is pinned to source commit
12ea99b2a57a354f2b7d6792f62a3d9313192fa7.
Evaluation protocol
Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ,
HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.YOUR-PHONE-NUMBER-IS-A-BEARER-TOKEN-CVID-CAPABILITY-BASED-REACHABILITY
Your Phone Number Is a Bearer Token
CVID — Capability-Validated Inbound Descriptors. Reachability as consumable,
revocable, cryptographically bound authority instead of persistent exposure.
Possession of an identifier is not permission to reach.
Abstract
A basic architectural weakness remains embedded in modern telecommunications: knowing how to reach someone is often practically equivalent to having permission to attempt contact.
A phone number, email address… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/YOUR-PHONE-NUMBER-IS-A-BEARER-TOKEN-CVID-CAPABILITY-BASED-REACHABILITY.benchability-fig4-capability-guided
BenchAbility Figure 4 -- capability_guided
One of two training mixtures drawn from the same frozen 884,143-row candidate pool, with the
same budget (60,000 intervention + 15,000 shared replay) and the same hyperparameters. The two
differ only in how the samples are chosen, which is the whole experiment.
arm
capability_guided
selection
by capability, gap-weighted from the Figure 3 scores
intervention rows
59,999
replay rows
15,000
shards
38
pool
884,143… See the full description on the dataset page: https://huggingface.co/datasets/realzL/benchability-fig4-capability-guided.CAPability
Dataset Card for Dataset Name
Visual caption benchmark Repo: CAPability
[🍎 Project Page] [📖 ArXiv Paper] [🧑💻 Github Repo] [🏆 Leaderboard]
Dataset Details
Visual captioning benchmarks have become outdated with the emergence of modern MLLMs, as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword extraction or object-centric evaluation… See the full description on the dataset page: https://huggingface.co/datasets/lntzm/CAPability.typhoon-s-sovereign-capability-dataset
Typhoon-S Training Assets
Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project.
Datasets
NitiBench (Legal Domain)
nitibench_train_rl.parquet - RL training set (8,211 examples)
nitibench_train_pretrain.parquet - Pretrain set (3,648 examples)
nitibench_train_sft.parquet - SFT set (3,648 examples)
nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-sovereign-capability-dataset.typhoon-s-sovereign-capability-dataset
Typhoon-S Training Assets
Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project.
Datasets
NitiBench (Legal Domain)
nitibench_train_rl.parquet - RL training set (8,211 examples)
nitibench_train_pretrain.parquet - Pretrain set (3,648 examples)
nitibench_train_sft.parquet - SFT set (3,648 examples)
nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-sovereign-capability-dataset.nanochat-depo-capability-data
Nanochat Depo Capability Pilot
This dataset is a deterministic natural-language rendering of the Depo directed-cycle
successor task. Each row contains shuffled operational records, one exact multi-hop
question, and its answer. Latent worlds are generated programmatically; no rows were
written or labeled by a language model.
Splits
Split
Worlds
Queries per world
Rows
Renderer family
train
32,768
4
131,072
incident handoff, six structural styles… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-depo-capability-data.nanochat-brevo-capability-data-10x
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for the complete recursive
prerequisite closure in a valid leaf-first order. Training uses project-planning
language; validation uses evidence synthesis; test uses build manifests. Eleven
deterministic structural styles vary wording, layout, and record order.
The latent graph generator and exact validator label every row. No language model
generated or labeled the data. Alternative valid orders… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-10x.milp_capability_3.2AutoBencher-capability.jsonai-capability-hiding-drift-detection-v0.1
What this dataset is
This dataset detects drift in capability-hiding patterns over time.
It compares:
baseline probe capability vs baseline expressed performance
current probe capability vs current expressed performance
whether a change in monitoring context explains a new gap
The goal is not blame.
The goal is early warning that oversight changes expression.
What it tests
You detect when the monitored/unmonitored gap:
newly appears
widens
changes shape
You also avoid… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-capability-hiding-drift-detection-v0.1.openrouter-model-pricing-capability-matrix
OpenRouter Model Pricing and Capability Matrix
This dataset packages 429 public OpenRouter model records across 75 developers into one analysis-ready table for AI procurement, model-selection workflows, router strategy, pricing research, and internal model-catalog dashboards.
It starts from OpenRouter's public models API and count endpoint, then enriches each model with its per-model endpoint metadata so buyers can compare token pricing, context windows, provider diversity… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/openrouter-model-pricing-capability-matrix.nanochat-capability-hardened-v5-20260714
Nanochat Capability Mixture
This corpus mixes depo, brevo, world_state, tool_routing with deterministic documents balancing. Training and validation rows only are included; source test rows remain excluded. Exact source manifests are retained under provenance/, and source-aware validation can replay every selected row against the authoritative datasets.
nanochat-brevo-capability-data
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for the complete recursive
prerequisite closure in a valid leaf-first order. Training uses project-planning
language; validation uses evidence synthesis; test uses build manifests. Six
deterministic structural styles vary wording, layout, and record order.
The latent graph generator and exact validator label every row. No language model
generated or labeled the data. Alternative valid orders are… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data.milp_capability_3.2_llamamilp_capability_3.2_gpt_3milp_capability_6.2qwen38-frontier-g01-capabilityai-coding-agent-pricing-and-capability-dataset
AI Coding Agent Pricing and Capability Dataset
A source-backed market-intelligence dataset for comparing AI coding agents and developer workflow agents across pricing, workflow support, release signals, repository activity, integrations, and public capability claims.
Each row represents one observed market signal tied to an official product page, official documentation page, official pricing page, public GitHub repository, or public GitHub release note. The dataset is built for… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/ai-coding-agent-pricing-and-capability-dataset.nanochat-four-capability-mixture-v3-5x-20260714-r1
Nanochat Capability Mixture
This corpus mixes depo, brevo, world_state, tool_routing with deterministic documents balancing. Training and validation rows only are included; source test rows remain excluded. Exact source manifests are retained under provenance/, and source-aware validation can replay every selected row against the authoritative datasets.
nanochat-capability-v4-mixed-20260714
Nanochat Capability Mixture
This corpus mixes depo, brevo, world_state, tool_routing with deterministic documents balancing. Training and validation rows only are included; source test rows remain excluded. Exact source manifests are retained under provenance/, and source-aware validation can replay every selected row against the authoritative datasets.
nanochat-brevo-capability-v4-49k-20260714
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for complete recursive
prerequisite closures in valid leaf-first orders. Each compact training document
reuses one graph for 4 worked questions, increasing answer
supervision without repeating the graph. Training balances
project_plan, build_manifest language; validation uses held-out
evidence synthesis; test uses build manifests. Eleven deterministic
structural styles vary wording, layout, and… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-v4-49k-20260714.nanochat-brevo-capability-data-v2
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for complete recursive
prerequisite closures in valid leaf-first orders. Each compact training document
reuses one graph for 4 worked questions, increasing answer
supervision without repeating the graph. Training uses project-planning language;
validation uses evidence synthesis; test uses build manifests. Eleven deterministic
structural styles vary wording, layout, and record order. Per-world… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-v2.milp_capability_6.2_labeledverify_conversational_history_capabilitynanochat-corrected-four-capability-mix-20260716-r2
Nanochat Capability Mixture
This corpus mixes depo_composition, brevo, world_state, tool_routing with deterministic documents balancing. Training and validation rows only are included; source test rows remain excluded. Exact source manifests are retained under provenance/, and source-aware validation can replay every selected row against the authoritative datasets.
europe-owid-population-covered-by-mobile-network-by-network-capability
Population Covered By Mobile Network By Network Capability | Europe (Our World in Data)
🇪🇺 915 observations · 42 Europe countries · 2000–2023 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 915 observations of Population Covered By Mobile Network By Network Capability data across 42 Europe countries, spanning 2000–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-population-covered-by-mobile-network-by-network-capability.nanochat-depo-brevo-capability-mixture
Nanochat Depo + Brevo Capability Mixture
This corpus deterministically interleaves Depo and Brevo documents and balances the two tasks by UTF-8 training-text bytes. Source manifests remain authoritative; test examples are excluded and capability evaluation is performed separately.
defendable-pain-federal-capability-statement-v0.1
Capability Statement Pain Receipt
"the one-pager" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 1 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
1 pain… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-federal-capability-statement-v0.1.capability-vectors
capability-vectors — PASS / FAIL contrast trajectories
This dataset bundles the exact PASS (comply) and FAIL (refuse) trajectories used to compute every direction in the
AlexWortega/capability-vectors abliteration-experiments repo.
Base model under study: AlexWortega/qwen35-4b-soyuz
(LoRA on Qwen3.5-4B, Soyuz-sft).
Each row is a single evaluation trajectory rendered to a chat-templated text blob, with the rollout's
source bench, task_id, and either a numeric reward (comply) or a… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/capability-vectors.
