datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qiskit-calibration-drift
Qiskit Calibration Drift
Calibration parameters from IBM Quantum Heron processors joined to ambient and space-weather conditions at the time of each measurement. Designed for time-series forecasting of qubit drift and for studying environmental coupling to superconducting calibrations.
A GitHub Action (source) polls backend.properties() on every available Heron device every 30 minutes and appends new calibration events keyed on (backend, property, qubit_a, qubit_b… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/qiskit-calibration-drift.GLM-5.3-Flash-calibration-activations-v1
GLM-5.3-Flash calibration activations v1 (BF16, natural routing)
Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048
tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in
and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up
input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth).
Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.calibration-scorecards
Prediction Market Calibration Scorecards
Monthly Brier + log-loss calibration breakdowns for Kalshi + Polymarket. Each month provides mean Brier, mean log-loss, per-venue and per-category breakdowns, and a 10-bucket calibration histogram (actual vs predicted). Published with a 14-day delay after month-end to capture late resolutions.
License and Use
This dataset is released under Creative Commons Attribution 4.0 International
(CC-BY-4.0;… See the full description on the dataset page: https://huggingface.co/datasets/SimpleFunctions/calibration-scorecards.pa-warm-start-sft-xl-calibrationInkling-Small-Multimodal-Calibration
Inkling-Small Multimodal Calibration
The exact 1,663 samples used for BF16 routed-expert importance collection
for Inkling-Small Mixed Quant GGUF.
This is calibration material, not a held-out evaluation benchmark.
The primary balanced pass is:
Category
Samples
Valid decoder tokens
Share
Text / reasoning
462
471,858
44.976%
Code / tool-oriented source text
205
209,715
19.989%
Real image / document
486
262,476
25.018%
Real speech audio
309
105,080
10.016%
Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.chai1-calibration-runDeepSeek-R1-Distill-Qwen-1.5B-Self-CalibrationThis dataset contains data for the paper Efficient Test-Time Scaling via Self-Calibration.
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model uncertainty on different tasks. For example, we can incorporate the model’s confidence into self-consistency by assigning each sampled response $y_i$ a confidence score $c_i$. Instead of treating all responses… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/DeepSeek-R1-Distill-Qwen-1.5B-Self-Calibration.ptq-calibration-corpus
PTQ Calibration Corpus
A compact, high-density corpus built specifically for post-training
quantization of chat, coding, and agentic language models.
This is not random web text and it is not a repackaged benchmark. The corpus
was deliberately assembled to exercise varied scientific and technical prose,
numeric structure, multilingual code, tool schemas, long debugging
trajectories, CUDA/Triton vocabulary, correctness recovery, and
architecture-sensitive optimization.
This is… See the full description on the dataset page: https://huggingface.co/datasets/RESMP-DEV/ptq-calibration-corpus.Llama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration.
🔥 Updates
[2025-3-3]: We released our paper.
[2025-2-25]: We released our codes, models and datasets.
🏴 Overview
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.Qwen2.5-7B-Instruct-Self-Calibration
Efficient Test-Time Scaling via Self-Calibration
This repository contains datasets used in the paper Efficient Test-Time Scaling via Self-Calibration. The datasets are used to evaluate the effectiveness of test-time scaling methods for LLMs. Each config_name in the metadata refers to a different reasoning dataset. More detailed descriptions of each dataset are needed. Consider adding a section for each config_name with a description, statistics, and any other relevant… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Qwen2.5-7B-Instruct-Self-Calibration.llm-forecast-calibration
LLM Forecast Calibration Study — GLM-5.3 on resolved Manifold Markets questions
Raw generation data for the study "Does sampling K times beat thinking harder?
A controlled study of LLM forecast calibration on resolved binary questions."
Source repo: EzraStone/llm-forecast-calibration.
Data mirrored from GitHub commit 0f12f71a2c2ec8c54cafeb4231fecb87e705e660.
All eight JSONL files match the source data byte for byte. The source repository
remains canonical for analysis code… See the full description on the dataset page: https://huggingface.co/datasets/ezra77/llm-forecast-calibration.eval-llama-3_1-8b-cpt-calibration
CPT calibration eval results
Per-window memorization eval results for every model state we trained in the
CPT-injection calibration experiment. Each subdirectory is one (model, dataset)
pair; results.parquet has one row per song with the standard cleanslate
sliding-window schema (coverage, mean_p_z, windows, content_id, ...).
See docs/calibration_experiment/ in the CleanSlate codebase for the full
experiment description.
Naming convention
eval_run10 — Llama-CPT… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/eval-llama-3_1-8b-cpt-calibration.jev-calibration-statistics
Confidence statistics for Jev and a self-judging Gemma 4 E2B
Aggregate statistics on the confidence scores from two judges in a retrieval benchmark: TypeSafe's Jev, pinned to jev-1.13.0, and Gemma 4 E2B judging its own work.
End to end, the pipeline with Jev making every decision did not beat the same pipeline with no judge: it scored 0.612 against 0.740, missed its main pre-registered bar, made about the same number of mistakes on questions both answered, and lost because it… See the full description on the dataset page: https://huggingface.co/datasets/clduab11/jev-calibration-statistics.less-is-moe-s1-calibration-128-seq8192
Less-is-MoE S1K calibration data — 128 samples, seq_length 8192
This is the fixed calibration artifact used to prune GPT-OSS-120B and
Qwen3.5-122B-A10B. It uses the same 128 source rows as the full-length variant:
yentinglin/s1K-1.1-trl-format revision
58a01564d278477da20ead1bcf1cde8e31f36251, train, followed by
Dataset.shuffle(seed=1234) and the first 128 nonempty messages rows.
For pruning, concatenate messages[].content with one space, tokenize with the
model's tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128-seq8192.test_calibration1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 5,
"total_frames": 8212,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mthirumalai/test_calibration1.less-is-moe-s1-calibration-128
Less-is-MoE S1K calibration data — 128 full-length samples
This repository contains the exact 128 S1K rows selected for Less-is-MoE
full-model pruning. The selection reproduces the released loader:
source: yentinglin/s1K-1.1-trl-format
revision: 58a01564d278477da20ead1bcf1cde8e31f36251
split: train
order: Dataset.shuffle(seed=1234)
samples: first 128 nonempty messages rows
sequence-length limit: none
truncation: disabled
padding: disabled
calibration.jsonl stores every… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128.summeval-annotated-latest
SummEval-LLMEval Dataset
Overview
The original SummEval dataset (Fabbri et al., 2021) consists of 1,600 summaries annotated by human expert evaluators using a 5-point Likert scale across 4 criteria: coherence, consistency, fluency, and relevance. These 1,600 summaries are based on 100 source articles from the CNN/DailyMail dataset (Hermann et al., 2015). For each source article, SummEval collects 16 summaries generated by 16 different automatic summarization systems. Each… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/summeval-annotated-latest.lecal_calibrationpol-dataset-text-no-url-calibration
Dataset Card for "pol-dataset-text-no-url-calibration"
More Information needed
eval_eval_calibration_test_calibratedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 294,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hrhraj/eval_eval_calibration_test_calibrated.qwen3.5-moe-awq-calibration
Qwen3.5 MoE AWQ Calibration Dataset
Calibration dataset for AWQ (Activation-Aware Weight Quantization) of
Qwen/Qwen3.5-35B-A3B and
Qwen/Qwen3.5-35B-A3B-Base.
Designed for MoE expert routing diversity: Qwen3.5-35B-A3B has 256 experts with 8
active per token, so calibration data needs broad domain coverage to exercise as many
routing paths as possible.
Sampling methodology
Source: PleIAs/common_corpus
(open multi-domain corpus with labeled collections)
Filtering:
Token… See the full description on the dataset page: https://huggingface.co/datasets/Lambent/qwen3.5-moe-awq-calibration.mtbench-annotated-latest
MT-Bench-Select Dataset
Introduction
The MT-Bench-Select dataset is a refined subset of the original MT-Bench dataset introduced by Zheng et al. (2023). The original MT-Bench dataset comprises 80 questions with answers generated by six models. Each question and each pair of models form an evaluation task, resulting in 1,200 tasks.
For this dataset, we used a curated subset of the original MT-Bench dataset, as prepared by the authors of the LLMBar paper (Zeng et al.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/mtbench-annotated-latest.eval_calibration_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 294,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hrhraj/eval_calibration_test.llmbar-annotated-latest
LLMBar-Select Dataset
Introduction
The LLMBar-Select dataset is a curated subset of the original LLMBar dataset introduced by Zeng et al. (2024). The LLMBar dataset consists of 419 instances, each containing an instruction paired with two outputs: one that faithfully follows the instruction and another that deviates while presenting superficially appealing qualities. It is designed to evaluate LLM-based evaluators more rigorously and objectively than previous benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/llmbar-annotated-latest.calibration-datasetlfm2.5-calibration-pack
Calibration Pack v0 — 5,000,000 Tokens Dataset
Mô hình mục tiêu: LFM2.5-2.6BQuy mô: 5,000,000 tokens (13,975 sequences)Ngày tạo: 2026-09-12 16:25:56
1. Cơ cấu Phân vùng (Partitions)
Phân vùng
Tệp Parquet
Mẫu (Seqs)
Số Tokens
Dung lượng
Short Language / General
partitions/01_short_language_1.5m.parquet
2,159
1,500,000
5.05 MB
Reasoning & Arithmetic
partitions/02_reasoning_1.0m.parquet
8,433
1,000,000
2.14 MB
Targeted Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/Neeze/lfm2.5-calibration-pack.eval_calibration_test_25514This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 592,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hrhraj/eval_calibration_test_25514.soarm101-test_calibration3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 445,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ricklon/soarm101-test_calibration3.CalibrationBench
LLM-as-a-Fuser JudgeBench results
This is the public, results-only dataset release for “Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution” (arXiv:2508.06225v3, DOI). The source code is at hzh030430/CalibrationBench, and the dataset repository is Han03430/CalibrationBench.
The package contains a 22,050-row core result release and a separate 350-row supplementary result table. The core consists of 9 Self-Confidence (SC) runs, 9 Multiple-Prompting (MP) runs… See the full description on the dataset page: https://huggingface.co/datasets/Han03430/CalibrationBench.qiskit-calibration-drift
IBM Quantum Calibration Drift Dataset
Continuously-updated calibration data from IBM Quantum hardware with concurrent environmental measurements. Enables correlation analysis between qubit performance and atmospheric/space weather conditions.
Overview
Property
Value
Update frequency
Every 30 minutes
Backends
ibm_fez (156 qubits), ibm_torino (133 qubits), ibm_marrakesh (156 qubits)
Total qubits
445
Collection method
Automated polling via GitHub Actions… See the full description on the dataset page: https://huggingface.co/datasets/nazimari/qiskit-calibration-drift.
