CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01phanerozoic /qiskit-calibration-drift Qiskit Calibration Drift Calibration parameters from IBM Quantum Heron processors joined to ambient and space-weather conditions at the time of each measurement. Designed for time-series forecasting of qubit drift and for studying environmental coupling to superconducting calibrations. A GitHub Action (source) polls backend.properties() on every available Heron device every 30 minutes and appends new calibration events keyed on (backend, property, qubit_a, qubit_b… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/qiskit-calibration-drift.tabulartime-series-forecasting10M<n<100M4 likes3.9k downloads14h agoHugging Face02malaiwah /GLM-5.3-Flash-calibration-activations-v1 GLM-5.3-Flash calibration activations v1 (BF16, natural routing) Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048 tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth). Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.tabularn<1K0 likes891 downloads24d agoHugging Face03SimpleFunctions /calibration-scorecards Prediction Market Calibration Scorecards Monthly Brier + log-loss calibration breakdowns for Kalshi + Polymarket. Each month provides mean Brier, mean log-loss, per-venue and per-category breakdowns, and a 10-bucket calibration histogram (actual vs predicted). Published with a 14-day delay after month-end to capture late resolutions. License and Use This dataset is released under Creative Commons Attribution 4.0 International (CC-BY-4.0;… See the full description on the dataset page: https://huggingface.co/datasets/SimpleFunctions/calibration-scorecards.tabularn<1K0 likes791 downloads21h agoHugging Face04geodesic-research /pa-warm-start-sft-xl-calibrationtabular100K<n<1M0 likes697 downloads10d agoHugging Face05Baekpica /Inkling-Small-Multimodal-Calibration Inkling-Small Multimodal Calibration The exact 1,663 samples used for BF16 routed-expert importance collection for Inkling-Small Mixed Quant GGUF. This is calibration material, not a held-out evaluation benchmark. The primary balanced pass is: Category Samples Valid decoder tokens Share Text / reasoning 462 471,858 44.976% Code / tool-oriented source text 205 209,715 19.989% Real image / document 486 262,476 25.018% Real speech audio 309 105,080 10.016% Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.tabulartext-generation1K<n<10K0 likes404 downloads14d agoHugging Face06swizman /chai1-calibration-rundocumentn<1K0 likes227 downloads2mo agoHugging Face07HINT-lab /DeepSeek-R1-Distill-Qwen-1.5B-Self-CalibrationThis dataset contains data for the paper Efficient Test-Time Scaling via Self-Calibration. We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model uncertainty on different tasks. For example, we can incorporate the model’s confidence into self-consistency by assigning each sampled response $y_i$ a confidence score $c_i$. Instead of treating all responses… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/DeepSeek-R1-Distill-Qwen-1.5B-Self-Calibration.tabularquestion-answering100K<n<1M0 likes169 downloads2y agoHugging Face08RESMP-DEV /ptq-calibration-corpus PTQ Calibration Corpus A compact, high-density corpus built specifically for post-training quantization of chat, coding, and agentic language models. This is not random web text and it is not a repackaged benchmark. The corpus was deliberately assembled to exercise varied scientific and technical prose, numeric structure, multilingual code, tool schemas, long debugging trajectories, CUDA/Triton vocabulary, correctness recovery, and architecture-sensitive optimization. This is… See the full description on the dataset page: https://huggingface.co/datasets/RESMP-DEV/ptq-calibration-corpus.tabulartext-generation1K<n<10K0 likes142 downloads1mo agoHugging Face09HINT-lab /Llama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration. 🔥 Updates [2025-3-3]: We released our paper. [2025-2-25]: We released our codes, models and datasets. 🏴󠁶󠁵󠁭󠁡󠁰󠁿 Overview We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.tabularquestion-answering100K<n<1M0 likes130 downloads2y agoHugging Face10HINT-lab /Qwen2.5-7B-Instruct-Self-Calibration Efficient Test-Time Scaling via Self-Calibration This repository contains datasets used in the paper Efficient Test-Time Scaling via Self-Calibration. The datasets are used to evaluate the effectiveness of test-time scaling methods for LLMs. Each config_name in the metadata refers to a different reasoning dataset. More detailed descriptions of each dataset are needed. Consider adding a section for each config_name with a description, statistics, and any other relevant… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Qwen2.5-7B-Instruct-Self-Calibration.tabulartext-generation100K<n<1M0 likes111 downloads2y agoHugging Face11ezra77 /llm-forecast-calibration LLM Forecast Calibration Study — GLM-5.3 on resolved Manifold Markets questions Raw generation data for the study "Does sampling K times beat thinking harder? A controlled study of LLM forecast calibration on resolved binary questions." Source repo: EzraStone/llm-forecast-calibration. Data mirrored from GitHub commit 0f12f71a2c2ec8c54cafeb4231fecb87e705e660. All eight JSONL files match the source data byte for byte. The source repository remains canonical for analysis code… See the full description on the dataset page: https://huggingface.co/datasets/ezra77/llm-forecast-calibration.tabularquestion-answering1K<n<10K0 likes93 downloads7d agoHugging Face12arpandeepk /eval-llama-3_1-8b-cpt-calibration CPT calibration eval results Per-window memorization eval results for every model state we trained in the CPT-injection calibration experiment. Each subdirectory is one (model, dataset) pair; results.parquet has one row per song with the standard cleanslate sliding-window schema (coverage, mean_p_z, windows, content_id, ...). See docs/calibration_experiment/ in the CleanSlate codebase for the full experiment description. Naming convention eval_run10 — Llama-CPT… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/eval-llama-3_1-8b-cpt-calibration.tabular1K<n<10K0 likes83 downloads5mo agoHugging Face13clduab11 /jev-calibration-statistics Confidence statistics for Jev and a self-judging Gemma 4 E2B Aggregate statistics on the confidence scores from two judges in a retrieval benchmark: TypeSafe's Jev, pinned to jev-1.13.0, and Gemma 4 E2B judging its own work. End to end, the pipeline with Jev making every decision did not beat the same pipeline with no judge: it scored 0.612 against 0.740, missed its main pre-registered bar, made about the same number of mistakes on questions both answered, and lost because it… See the full description on the dataset page: https://huggingface.co/datasets/clduab11/jev-calibration-statistics.tabularn<1K0 likes62 downloads1d agoHugging Face14jayzou3773 /less-is-moe-s1-calibration-128-seq8192 Less-is-MoE S1K calibration data — 128 samples, seq_length 8192 This is the fixed calibration artifact used to prune GPT-OSS-120B and Qwen3.5-122B-A10B. It uses the same 128 source rows as the full-length variant: yentinglin/s1K-1.1-trl-format revision 58a01564d278477da20ead1bcf1cde8e31f36251, train, followed by Dataset.shuffle(seed=1234) and the first 128 nonempty messages rows. For pruning, concatenate messages[].content with one space, tokenize with the model's tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128-seq8192.tabulartext-generationn<1K0 likes60 downloads4d agoHugging Face15mthirumalai /test_calibration1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 5, "total_frames": 8212, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mthirumalai/test_calibration1.tabularrobotics1K<n<10K0 likes59 downloads7mo agoHugging Face16jayzou3773 /less-is-moe-s1-calibration-128 Less-is-MoE S1K calibration data — 128 full-length samples This repository contains the exact 128 S1K rows selected for Less-is-MoE full-model pruning. The selection reproduces the released loader: source: yentinglin/s1K-1.1-trl-format revision: 58a01564d278477da20ead1bcf1cde8e31f36251 split: train order: Dataset.shuffle(seed=1234) samples: first 128 nonempty messages rows sequence-length limit: none truncation: disabled padding: disabled calibration.jsonl stores every… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128.tabulartext-generationn<1K0 likes55 downloads4d agoHugging Face17bay-calibration-llm-evaluators /summeval-annotated-latest SummEval-LLMEval Dataset Overview The original SummEval dataset (Fabbri et al., 2021) consists of 1,600 summaries annotated by human expert evaluators using a 5-point Likert scale across 4 criteria: coherence, consistency, fluency, and relevance. These 1,600 summaries are based on 100 source articles from the CNN/DailyMail dataset (Hermann et al., 2015). For each source article, SummEval collects 16 summaries generated by 16 different automatic summarization systems. Each… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/summeval-annotated-latest.tabular1K<n<10K0 likes51 downloads2y agoHugging Face18Toby0614 /lecal_calibrationtabular10K<n<100K0 likes48 downloads6mo agoHugging Face19BrennanGambling /pol-dataset-text-no-url-calibration Dataset Card for "pol-dataset-text-no-url-calibration" More Information needed tabular100K<n<1M0 likes42 downloads3y agoHugging Face20hrhraj /eval_eval_calibration_test_calibratedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 294, "total_tasks": 1, "total_videos": 1, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hrhraj/eval_eval_calibration_test_calibrated.tabularroboticsn<1K0 likes41 downloads1y agoHugging Face21Lambent /qwen3.5-moe-awq-calibration Qwen3.5 MoE AWQ Calibration Dataset Calibration dataset for AWQ (Activation-Aware Weight Quantization) of Qwen/Qwen3.5-35B-A3B and Qwen/Qwen3.5-35B-A3B-Base. Designed for MoE expert routing diversity: Qwen3.5-35B-A3B has 256 experts with 8 active per token, so calibration data needs broad domain coverage to exercise as many routing paths as possible. Sampling methodology Source: PleIAs/common_corpus (open multi-domain corpus with labeled collections) Filtering: Token… See the full description on the dataset page: https://huggingface.co/datasets/Lambent/qwen3.5-moe-awq-calibration.tabulartext-generationn<1K0 likes39 downloads7mo agoHugging Face22bay-calibration-llm-evaluators /mtbench-annotated-latest MT-Bench-Select Dataset Introduction The MT-Bench-Select dataset is a refined subset of the original MT-Bench dataset introduced by Zheng et al. (2023). The original MT-Bench dataset comprises 80 questions with answers generated by six models. Each question and each pair of models form an evaluation task, resulting in 1,200 tasks. For this dataset, we used a curated subset of the original MT-Bench dataset, as prepared by the authors of the LLMBar paper (Zeng et al.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/mtbench-annotated-latest.tabular1K<n<10K0 likes34 downloads2y agoHugging Face23hrhraj /eval_calibration_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 294, "total_tasks": 1, "total_videos": 1, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hrhraj/eval_calibration_test.tabularroboticsn<1K0 likes34 downloads1y agoHugging Face24bay-calibration-llm-evaluators /llmbar-annotated-latest LLMBar-Select Dataset Introduction The LLMBar-Select dataset is a curated subset of the original LLMBar dataset introduced by Zeng et al. (2024). The LLMBar dataset consists of 419 instances, each containing an instruction paired with two outputs: one that faithfully follows the instruction and another that deviates while presenting superficially appealing qualities. It is designed to evaluate LLM-based evaluators more rigorously and objectively than previous benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/llmbar-annotated-latest.tabular1K<n<10K0 likes33 downloads2y agoHugging Face25szymonrucinski /calibration-datasettabular1K<n<10K0 likes32 downloads2y agoHugging Face26Neeze /lfm2.5-calibration-pack Calibration Pack v0 — 5,000,000 Tokens Dataset Mô hình mục tiêu: LFM2.5-2.6BQuy mô: 5,000,000 tokens (13,975 sequences)Ngày tạo: 2026-09-12 16:25:56 1. Cơ cấu Phân vùng (Partitions) Phân vùng Tệp Parquet Mẫu (Seqs) Số Tokens Dung lượng Short Language / General partitions/01_short_language_1.5m.parquet 2,159 1,500,000 5.05 MB Reasoning & Arithmetic partitions/02_reasoning_1.0m.parquet 8,433 1,000,000 2.14 MB Targeted Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/Neeze/lfm2.5-calibration-pack.tabular10K<n<100K0 likes32 downloads11d agoHugging Face27hrhraj /eval_calibration_test_25514This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 592, "total_tasks": 1, "total_videos": 1, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hrhraj/eval_calibration_test_25514.tabularroboticsn<1K0 likes29 downloads1y agoHugging Face28ricklon /soarm101-test_calibration3This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 1, "total_frames": 445, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ricklon/soarm101-test_calibration3.tabularroboticsn<1K0 likes26 downloads6mo agoHugging Face29Han03430 /CalibrationBench LLM-as-a-Fuser JudgeBench results This is the public, results-only dataset release for “Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution” (arXiv:2508.06225v3, DOI). The source code is at hzh030430/CalibrationBench, and the dataset repository is Han03430/CalibrationBench. The package contains a 22,050-row core result release and a separate 350-row supplementary result table. The core consists of 9 Self-Confidence (SC) runs, 9 Multiple-Prompting (MP) runs… See the full description on the dataset page: https://huggingface.co/datasets/Han03430/CalibrationBench.tabulartext-classification10K<n<100K0 likes26 downloads2mo agoHugging Face30nazimari /qiskit-calibration-drift IBM Quantum Calibration Drift Dataset Continuously-updated calibration data from IBM Quantum hardware with concurrent environmental measurements. Enables correlation analysis between qubit performance and atmospheric/space weather conditions. Overview Property Value Update frequency Every 30 minutes Backends ibm_fez (156 qubits), ibm_torino (133 qubits), ibm_marrakesh (156 qubits) Total qubits 445 Collection method Automated polling via GitHub Actions… See the full description on the dataset page: https://huggingface.co/datasets/nazimari/qiskit-calibration-drift.tabulartime-series-forecasting100K<n<1M0 likes25 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.