datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
caliber-adaptive-toolcall-v2
CALIBER — Calibrated Adaptive-effort Live-tool Instruction corpus for Benchmarked, Execution-verified Reasoning
A tool-caller is best not when it thinks the most or holds the most instruments, but when it
spends deliberation only where deliberation pays — think hard on ambiguous / multi-step turns,
act directly on unambiguous single-tool turns, and abstain / ask / refuse correctly — and
every one of those decisions is execution/rubric-verified, decontaminated, and grounded in… See the full description on the dataset page: https://huggingface.co/datasets/Jables/caliber-adaptive-toolcall-v2.caliber-extension-gemma4-e2b-grpo-rollouts
CALIBER Extension — Gemma4-E2B GRPO Rollouts
Training rollouts from matched GRPO arms on google/gemma-4-E2B-it
(new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps).
Subsets
subset
arm
τ
prior
rows
mean reward_total
accuracy
full schema
caliber
vanilla CALIBER
0.0
—
1600
2.298
0.514
0.664
mink
Min-K% prior
1.0
mink_0.2
4800
2.506
0.520
0.680
minkpp
Min-K++% prior
1.0
minkpp_0.2
4800
2.637
0.541
0.726
Load:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.caliber-adaptive-toolcall-v1
CALIBER — Calibrated Adaptive-effort Live-tool Instruction corpus for Benchmarked, Execution-verified Reasoning
A tool-caller is best not when it thinks the most or holds the most instruments, but when it
spends deliberation only where deliberation pays — think hard on ambiguous / multi-step turns,
act directly on unambiguous single-tool turns, and abstain / ask / refuse correctly — and
every one of those decisions is execution/rubric-verified, decontaminated, and grounded in… See the full description on the dataset page: https://huggingface.co/datasets/Jables/caliber-adaptive-toolcall-v1.
