CoolFace
Datasetpublic

Bauxitiego/surface-code-syndromes

Surface code syndromes, ML-ready Real quantum error correction data from Google Quantum AI's Sycamore processor, reshaped so you can train a model on it without knowing what a detector error model is. This is a reformatting, not new data. The measurements are Google's, released under CC-BY-4.0 alongside their 2023 Nature paper. What is added here is structure: a flat schema, fixed splits, reshaping metadata, and the published decoder predictions bundled per shot. Why… See the full description on the dataset page: https://huggingface.co/datasets/Bauxitiego/surface-code-syndromes.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes119downloads
Dataset Card

Surface code syndromes, ML-ready

Real quantum error correction data from Google Quantum AI's Sycamore processor, reshaped so you can train a model on it without knowing what a detector error model is.

This is a reformatting, not new data. The measurements are Google's, released under CC-BY-4.0 alongside their 2023 Nature paper. What is added here is structure: a flat schema, fixed splits, reshaping metadata, and the published decoder predictions bundled per shot.

Why this exists

The original release is excellent and nearly unusable from a machine learning workflow. It ships as bit-packed .b8 files across 130 directories keyed by a naming convention, and recovering the time-series structure means parsing the accompanying stim circuits to read detector coordinates. You need to understand quantum error correction before you can look at a single label.

Decoding a surface code is, stripped of physics, binary sequence classification on sparse binary channels with exact labels. That should be accessible to anyone who trains models. This dataset makes it so.

The task

Each row is one experimental shot. The input is a syndrome: a multi-channel binary time series of parity-check measurements. The label is whether the logical qubit ended up flipped.

python
from datasets import load_dataset
import numpy as np

ds = load_dataset("Bauxitiego/surface-code-syndromes")
row = ds["train"][0]
syndrome = np.array(row["syndrome"], np.uint8).reshape(row["time_steps"], row["width"])
label = row["label"]

Fields

FieldTypeDescription
experimentstringSource directory name, e.g. surface_code_bZ_d5_r25_center_5_5
basisstringMemory basis, X or Z
distanceint8Code distance, 3 or 5
roundsint16Syndrome extraction rounds, odd values 1 to 25
time_stepsint8First dimension of the reshaped syndrome, always rounds + 1
widthint16Second dimension, the widest round
round_widthslist[uint8]Real detectors per time step, for reconstructing the mask
syndromelist[uint8]Flattened [time_steps, width], zero-padded
labeluint8Ground truth: did the logical observable flip
pymatchinguint8Minimum-weight perfect matching prediction
correlated_matchinguint8Correlated matching prediction
tensor_networkuint8Tensor network contraction prediction, null where not run

Padding matters

Rounds are ragged. The first time step carries only the stabilisers of the memory basis and the last comes from the data qubit measurements, so both are half the width of the middle rounds. A distance-5, 25-round experiment has widths [12, 24, 24, ..., 24, 12].

Rows are right-padded to the widest round, so some slots are structurally empty. Use round_widths to build a mask and make sure padding cannot reach your model's output:

python
mask = np.zeros((row["time_steps"], row["width"]), np.uint8)
for t, w in enumerate(row["round_widths"]):
    mask[t, :w] = 1

A padded zero and an observed no-detection are different things. Treating them as the same is a silent bug, and it will not show up as a training failure.

Baselines come with the data

Google ran three decoders and published their per-shot predictions, so you can compute the numbers you need to beat without installing a decoder:

python
import numpy as np
sub = ds["test"].filter(lambda r: r["distance"] == 5 and r["rounds"] == 25 and r["basis"] == "Z")
label = np.array(sub["label"])
for name in ("pymatching", "correlated_matching"):
    print(name, (np.array(sub[name]) != label).mean())

On distance 5, 25 rounds, basis Z, test split (10,000 shots):

DecoderLogical error rate
Always predict "no flip"0.5098
pymatching0.4399
Correlated matching0.4028

Those look high because 25 rounds accumulate error. The meaningful quantity is error per round, and the meaningful comparison is against these baselines rather than against zero.

Correlated matching and tensor network contraction are strong. Tensor network contraction is close to optimal for these circuits. Beating pymatching is a reasonable target; beating tensor network contraction is not, and a paper claiming to would need extraordinary evidence.

Splits

70 / 10 / 20 train, validation, test, stratified within each experiment and shuffled with a fixed seed (20260806) before splitting.

The shuffle is deliberate. Device shots arrive in acquisition order and calibration drifts over a run, so an unshuffled split would train on one period of the experiment and test on another, which measures drift rather than decoding.

SplitRows
train4,550,000
validation650,000
test1,300,000

130 experiments: two bases, distances 3 and 5, odd round counts 1 through 25, and for distance 3 four different positions on the chip.

Limitations

  • One device, one calibration window, mid-2022. Error rates on current hardware are lower.
  • Distance 3 and 5 only. The 2024 follow-up reached distance 7 below threshold; that data is a separate release and is not included here.
  • 50,000 shots per experiment. Ample for evaluation, thin for training large models.
  • No sweep over physical error rate. The device has the error rates it has.
  • The four distance-3 chip positions have genuinely different noise. Pooling them trains a decoder that is worse on each than a per-position decoder would be. That is a real effect worth measuring, not a defect.

Citation

Cite the original data. This dataset is a reformatting and claims no measurement credit.

bibtex
@article{google2023suppressing,
  title   = {Suppressing quantum errors by scaling a surface code logical qubit},
  author  = {{Google Quantum AI}},
  journal = {Nature},
  volume  = {614},
  pages   = {676--681},
  year    = {2023},
  doi     = {10.1038/s41586-022-05434-1}
}

Original data: zenodo.org/records/6804040, CC-BY-4.0.

License

CC-BY-4.0, inherited from the source release. Attribution to Google Quantum AI is required.

Related

Conversion code, training code, and a study of when learned decoders beat matching: github.com/Bauxitiego/qec-neural-decoder