datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantized-llama-3.1-leaderboard-v2-evals
Open LLM Leaderboard v2 Benchmark Results
This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models.
These evaluations were produced with lm-evaluation-harness by running the following command:
lm_eval \
--model vllm \
--model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \
--apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.quant-fidelity-registry
Quantization Fidelity Registry
A public, schema'd, receipt-backed, cross-model index of quantization quality measurements.
It exists to answer one question that nothing else answers today: show me every measured quant of
model X, with its fidelity number and enough provenance to know whether that number means anything.
It is the sibling of 0xSero/local-ai-registry,
which answers how fast, how much VRAM, how much money. This one answers how faithful. Ids and the
huggingface… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/quant-fidelity-registry.exams-basic-and-quantum-cryptography-and-security-latex
Open Problem Exams: Cryptography and Security (LaTeX)
A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions.
Dataset Overview
Institution
Files
Topics
Questions
Caltech & TU Delft
8
38
145
EPFL
6
19
86
ETH Zurich
1
14
37
MIT
3
33
79
Total
18
104
347
Difficulty Distribution
Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.szl-quant-sft-v1
szl-quant-sft-v1 — training rows with signed lineage
Every row in this dataset is derived deterministically from a DSSE-signed backtest receipt and is recomputable bit-exact from content-addressed archives. No row was hand-written, scraped, or synthesized by a model.
Lineage (verifiable end-to-end)
CoinGecko daily closes (REPORTED venue feed)
→ szl-quant MEASURED walk-forward backtests → DSSE-signed receipts… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-quant-sft-v1.QuantumChem_Testbank_3000
Hold-out test set for evaluating quantaMind.
Qiskit-QuantumKatas
Qiskit QuantumKatas
A benchmark dataset for evaluating Large Language Models on quantum computing code generation tasks using Qiskit.
Dataset Description
This dataset contains 350 quantum computing tasks translated from Microsoft's QuantumKatas (originally in Q#) to Qiskit (Python). It is designed for evaluating LLMs on their ability to generate correct quantum computing code.
Supported Tasks
Code Generation: Given a natural language description and function… See the full description on the dataset page: https://huggingface.co/datasets/Qiskit/Qiskit-QuantumKatas.quant_exploration
Examining LLM Quantization Impact
This document is a comparative analysis of qualitative performance degradation across Llama.cpp quantization within a single 2x7B model. My hope is that it will help people unfamiliar with quant impacts get a sense of how quantization will affect output.
Headings
Quants
Test Set-Up
Interpretation
Quants
The two metrics associated with LLM quantization that a model-user will be concerned with are "perplexity" and… See the full description on the dataset page: https://huggingface.co/datasets/christopherthompson81/quant_exploration.v41-quant-workerHPLT3_DE_0.8_Quantilequantqa
QuantQA: Quantitative Finance Interview Questions
QuantQA is a curated dataset of 519 interview questions sourced from leading quantitative trading firms including Jane Street, Citadel, Two Sigma, Optiver, and SIG, in collaboration with CoachQuant.
Topic Distribution
Topic
Coverage
Probability
67%
Combinatorics
22%
Expected Value
21%
Conditional Probability
14%
Game Theory
11%
Note: Questions may cover multiple topics
Training Results… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/quantqa.quantum-physics-0.6-corpus
quantum-physics-0.6-corpus
Dataset Description
This is a domain-specific corpus created using ontology-guided filtering from FineWeb-Edu.
Dataset Creation
Source: HuggingFaceFW/fineweb-edu
Filtering Method: Semantic similarity to subdomain centroids (embedding-based)
Pipeline: Ontology-Guided Domain Corpus Builder
Dataset Structure
Each chunk contains:
text: The text content (256-512 tokens)
subdomain_id: Assigned subdomain
similarity_score:… See the full description on the dataset page: https://huggingface.co/datasets/konsman/quantum-physics-0.6-corpus.gmat-quant-corpus
GMAT Quant Corpus for Solver + Retrieval
This dataset is intended for retrieval over GMAT-style quantitative teaching content.
Files
gmat_hf_chunks.jsonl — retrieval chunks used by the app
gmat_question_seed.jsonl — question seed data
gmat_topic_index.json — topic metadata/index
Default dataset viewer
The default dataset viewer is configured to load only:
gmat_hf_chunks.jsonl
This avoids schema conflicts with the other support files in the repository.
hemmingway-1-omlx-quantization-evidence-v2
Hemmingway-1 Quantization Evidence v2
This package records two local evidence lanes for the Hemmingway-1 oQ4e build: teacher-forced numerical fidelity against a BF16 reference, and controlled runtime telemetry on Apple Silicon. It complements the frozen blind-preference study in Hemmingway-1 oMLX Quantization Benchmark v1.
This dataset is sixstringzen/hemmingway-1-omlx-quantization-evidence-v2. The quality dataset remains unchanged because blind preference, distribution fidelity… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-evidence-v2.HPLT3_DE_0.9_QuantileHPLT3_DE_0.9_Quantile_Adult_Filtered_Propelatis-quantile-datasets-Qwen3-4B-Basequantum-physics-0.6
quantum-physics-0.6
Dataset Description
This is a domain-specific corpus created using ontology-guided filtering from FineWeb-Edu.
Dataset Creation
Source: HuggingFaceFW/fineweb-edu
Filtering Method: Semantic similarity to subdomain centroids (embedding-based)
Pipeline: Ontology-Guided Domain Corpus Builder
Dataset Structure
Each chunk contains:
text: The text content (256-512 tokens)
subdomain_id: Assigned subdomain… See the full description on the dataset page: https://huggingface.co/datasets/konsman/quantum-physics-0.6.hemmingway-1-omlx-quantization-benchmark-v1
Hemmingway-1 oMLX Quantization Benchmark
This is the public-safe benchmark package for the Hemmingway-1 oMLX
quantization study on Apple Silicon.
Altworld developed and published
Hemmingway-1. Bobby Pierce
published these quantizations and the evaluation package. The
collection
links the upstream model and all six builds.
Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching
across reversed packets. Read CORRECTION.md before using the
aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.glm53-fixture-0.1B-fidelity-quant-int4-v1
GLM-5.3-Flash-0.1B fixture — candidate fidelity dataset, toy RTN-int4 routed experts (hidden form)
The numbers in this dataset are meaningless as quantization quality.
The weights are random (inference-optimization/GLM-5.3-Flash-0.1B-A0.1B is an
architectural fixture), and the quantizer is deliberately crude. This exists so
that step 3 of the three-step fidelity architecture has two real datasets to
compare, and so that anyone can see what a candidate capture looks like
next to… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-fixture-0.1B-fidelity-quant-int4-v1.tis-quantile-datasets-Olmo-3-1025-7Bchess_puzzles_10k_in_pgn_san
Lichess Puzzle Database …
Lichess Puzzle Database ▸ Mate‑in‑1/2/3 Subset
A 10k slice of the official Lichess puzzle corpus.
Note: The board is updated to the position that arises after the opponent blunders.
Source
Original dataset: https://huggingface.co/datasets/Lichess/chess-puzzles
What’s Inside
Category
Count
Mate‑in‑1
3333
Mate‑in‑2
3333
Mate‑in‑3
3333
Total
9999
Format Improvements
Field
Old
New… See the full description on the dataset page: https://huggingface.co/datasets/quantum24/chess_puzzles_10k_in_pgn_san.quantitative-finance-reasoning
Dataset Description
Dataset Summary
This dataset contains question-answer pairs focused on quantitative finance, covering topics such as option pricing, stochastic calculus (Brownian motion, Itô's Lemma), probability theory, and financial modeling assumptions. Each instance includes a question, a detailed ground-truth solution (often resembling textbook explanations or interview answers), a multi-step reasoning trace generated by a Gemini Pro model, and a structured validation of… See the full description on the dataset page: https://huggingface.co/datasets/VedantPadwal/quantitative-finance-reasoning.opensre-incident-trajectories
OpenSRE Incident-Diagnosis Trajectories
Graded, multi-step SRE incident-diagnosis trajectories. A frozen LLM reads evidence through
diagnostic tools (describe_pod / get_events / get_logs / get_metrics / query_traces / …),
states a root cause + category + fix, and is scored on substance against ground truth. Built as a
HUD v6 RL environment with a deliberate model spanning set so difficulty is legible and the
within-group reward spread is real (the GRPO learning signal).
197… See the full description on the dataset page: https://huggingface.co/datasets/quantranger/opensre-incident-trajectories.tamperbench-quantization-qwen3-4b
TamperBench + Quantization: Does Compression Act as Implicit Tampering?
Motivation
TamperBench evaluates explicit tampering attacks (LoRA fine-tuning, jailbreak-tuning, etc.) on LLM safety guards. Catastrophic Failure of LLM Unlearning via Quantization shows that quantization can undo safety-trained behaviors.
This experiment bridges these two lines of work by adding quantization as a deployment-realistic perturbation to the TamperBench evaluation protocol. We… See the full description on the dataset page: https://huggingface.co/datasets/b0sungk1m/tamperbench-quantization-qwen3-4b.quantum-worldline-research
Quantum Worldline Research Data
Structured research data from the Quantum Worldline project - an AI-assisted research program investigating holographic forces in MERA tensor networks, worldline path integrals on AdS spacetime, and quantum simulation of lattice gauge theories.
Dataset Description
This dataset contains the complete structured output of the Quantum Worldline multi-agent research system, which automates the research cycle: discover - hypothesize - gate - test… See the full description on the dataset page: https://huggingface.co/datasets/Jonboy648/quantum-worldline-research.bbq
BBQ
Repository for the Bias Benchmark for QA dataset.
https://github.com/nyu-mll/BBQ
Authors: Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R. Bowman.
This repository is a fork of https://huggingface.co/datasets/heegyu/bbq, and adds the "All" configuration containing all subsets.
About BBQ (paper abstract)
It is well documented that NLP models learn social biases, but little work has been done… See the full description on the dataset page: https://huggingface.co/datasets/quantiles/bbq.quantum-circuits-21k
Quantum Circuits Dataset — v2 (21K)
A synthetic dataset of validated natural language → OpenQASM 2.0 circuit pairs for training quantum circuit generation models. To our knowledge the largest publicly available dataset of validated NL→QASM pairs specifically designed for generative model training.
Used to train the QuantumGPT-124M model series.
Quick Start
from datasets import load_dataset
# v2 training set (21K samples, recommended)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/merileijona/quantum-circuits-21k.quantization-cache-amplification
Quantization as Cache Amplification
Trillion-Parameter Mixture-of-Experts Inference on a Commodity Laptop
Kavin Kumar, Neural Metrics
📄 Read the paper — 11 pages
What this is
Weight quantization is usually justified as footprint reduction. This work argues
that for offloaded mixture-of-experts inference that framing misses the leverage.
The binding resource is not storage capacity but the fraction of expert slots
resident in DRAM — and storage traffic depends on… See the full description on the dataset page: https://huggingface.co/datasets/mkvn/quantization-cache-amplification.grok-1-ternary-quant-experiments
Grok-1 SAAQ quantization / route-preservation experiments
Dataset author: Raul Montoya Cardenas (rmems)
SAAQ stands for Spiking Adaptive Activity Quantization, a term coined by
the dataset author.
Attribution: Grok Build: Grok 4.5 (high) packaged the original 2026-08-10
dataset. Codex: GPT-5.6-Sol (OpenAI) implemented, executed, validated, and published the
canonical issue #85 v4 evidence added on 2026-08-24.
Personal research measuring route preservation when packing open… See the full description on the dataset page: https://huggingface.co/datasets/rmems/grok-1-ternary-quant-experiments.quantum-ood-benchmark
