| --- |
| license: mit |
| library_name: pytorch |
| tags: |
| - quantum-computing |
| - quantum-error-correction |
| - surface-code |
| - syndrome-decoding |
| - physics |
| datasets: |
| - Bauxitiego/surface-code-syndromes |
| --- |
| |
| # Neural decoders for the surface code |
|
|
| Weights from a study of when learned decoders beat minimum-weight perfect matching, and why. |
| Anyone can train a network to decode a surface code; that has been done since 2017. The |
| question here is narrower. Matching is near-optimal when its noise model is right, so the |
| only place a learned decoder can win is where that model is wrong. This measures how much it |
| wins by, and at what point it stops. |
|
|
| Three architectures, all within 1.2x of each other on parameter count and given identical |
| optimiser, schedule and step budget. Comparing a big model to a small one, or one tuned |
| harder than another, measures capacity or patience rather than architecture. |
|
|
| ## Using a checkpoint |
|
|
| Weights, the decision threshold picked on validation, and the padding mask all travel in the |
| same file. A checkpoint that makes you reconstruct the threshold by hand is not really |
| reusable. |
|
|
| ```python |
| import torch |
| ckpt = torch.load("mlp_A_d3_r3_p0.003.pt", weights_only=False) |
| print(ckpt["architecture"], ckpt["distance"], ckpt["p"], ckpt["logical_error_rate"]) |
| ``` |
|
|
| Model definitions live in [github.com/Bauxitiego/qec-neural-decoder](https://github.com/Bauxitiego/qec-neural-decoder). |
| Filenames encode architecture, regime, distance, rounds and physical error rate. |
|
|
| ## The result |
|
|
| Regime B holds total noise constant and varies only its structure: the base error rate is |
| scaled down by bisection until detection-event density matches the uniform control, so a |
| "correlated noise" arm cannot secretly be a "more noise" arm. Matching is run twice, once |
| with the true noise model and once with the uniform one it would actually have on hardware. |
|
|
| | d | noise | MWPM true | MWPM mis-spec | penalty | best neural | vs mis-spec | |
| |---|---|---|---|---|---|---| |
| | 3 | `s0_x0` | 0.00652 | 0.00652 | 1.00x | 0.00617 | indistinguishable | |
| | 3 | `s0_x0.0005` | 0.01531 | 0.02835 | 1.85x | 0.01536 | neural | |
| | 3 | `s1.75_x0` | 0.00639 | 0.00656 | 1.03x | 0.00591 | neural | |
| | 3 | `s1.75_x0.0015` | 0.01974 | 0.03949 | 2.00x | 0.02006 | neural | |
| | 3 | `s1_x0` | 0.00695 | 0.00702 | 1.01x | 0.00646 | neural | |
| | 3 | `s1_x0.0005` | 0.01229 | 0.02002 | 1.63x | 0.01195 | neural | |
| | 5 | `s0_x0` | 0.00334 | 0.00334 | 1.00x | 0.01111 | mwpm | |
| | 5 | `s0_x0.0005` | 0.01603 | 0.02194 | 1.37x | 0.03611 | mwpm | |
| | 5 | `s1.75_x0` | 0.00290 | 0.00335 | 1.16x | 0.01052 | mwpm | |
| | 5 | `s1.75_x0.0015` | 0.02515 | 0.03575 | 1.42x | 0.04158 | mwpm | |
| | 5 | `s1_x0` | 0.00317 | 0.00356 | 1.12x | 0.01507 | mwpm | |
| | 5 | `s1_x0.0005` | 0.01117 | 0.01439 | 1.29x | 0.02854 | mwpm | |
|
|
| At distance 3 the networks track true-model matching and beat the mis-specified version by up |
| to 2.00x. Two rows are worth reading closely. Per-qubit rate spread on its own barely |
| costs matching anything, around 1.01x, because a mis-weighted graph is still roughly the |
| right graph. Crosstalk costs it 1.85x, because correlated errors have no edge to live on. |
| Rate variation is survivable; correlation is not. |
|
|
| At distance 5 matching wins every row. The next section is why. |
|
|
| ## Why distance 5 loses |
|
|
| | training shots | epochs | logical error rate | vs MWPM | |
| |---|---|---|---| |
| | 120,000 | 8 | 0.04079 | 12.1x | |
| | 400,000 | 25 | 0.01814 | 5.4x | |
| | 800,000 | 40 | 0.01118 | 3.3x | |
|
|
| The sweep above ran at 400,000 shots and 12 epochs, between the first two rows. At distance 3 the networks saturate by the second row and settle slightly under matching. At distance 5 they are still improving at the largest budget tested, going 12.1x to 5.4x to 3.3x without flattening. |
|
|
| So the distance-5 numbers are a statement about how long I trained. Read them as an architecture result and you will draw the wrong conclusion, which is the whole reason this section is here. |
|
|
| ## Regime A: uniform noise |
|
|
| The control. Matching's model is exactly right here, so it should win, and mostly it does. |
|
|
| | d | p | trivial | MWPM | best neural | arch | winner | |
| |---|---|---|---|---|---|---| |
| | 3 | 0.001 | 0.02308 | 0.00072 | 0.00092 | mlp | indistinguishable | |
| | 3 | 0.003 | 0.06567 | 0.00641 | 0.00617 | mlp | indistinguishable | |
| | 3 | 0.005 | 0.10454 | 0.01668 | 0.01609 | mlp | indistinguishable | |
| | 3 | 0.01 | 0.18744 | 0.06056 | 0.05562 | mlp | neural | |
| | 5 | 0.001 | 0.05801 | 0.00010 | 0.00262 | gnn | mwpm | |
| | 5 | 0.003 | 0.15393 | 0.00338 | 0.01111 | gnn | mwpm | |
| | 5 | 0.005 | 0.22917 | 0.01401 | 0.03736 | gnn | mwpm | |
| | 5 | 0.01 | 0.35625 | 0.08121 | 0.15477 | gnn | mwpm | |
|
|
| One row is not a tie: at distance 3 and p=0.01 the MLP beats matching outright with |
| non-overlapping intervals. That was not the expected outcome in the regime designed to |
| favour matching. |
|
|
| The architecture ordering also flips with distance. The plain MLP is best at distance 3, the |
| graph network at distance 5. The graph prior costs more than it returns until the code is |
| large enough for the structure to carry information. |
|
|
| ## Real device: Google Sycamore |
|
|
| Syndromes from Google's 2023 experiment, with their own published decoder predictions as |
| baselines, scored on identical held-out shots. No reimplementation sits between these models |
| and the comparison. |
|
|
| | d | rounds | trivial | pymatching | correlated | best neural | vs pymatching | |
| |---|---|---|---|---|---|---| |
| | 3 | 5 | 0.29340 | 0.14790 | 0.13760 | 0.15720 | indistinguishable | |
| | 3 | 25 | 0.49010 | 0.43020 | 0.42410 | 0.48920 | mwpm | |
| | 5 | 5 | 0.38690 | 0.14730 | 0.12750 | 0.26430 | mwpm | |
|
|
| These lose, and the 25-round row loses badly enough to be worth stating plainly: 0.489 |
| against a trivial floor of 0.490 means the model learned essentially nothing. Each experiment |
| has 50,000 shots, so 35,000 to train on, against a 200-detector input with a base flip rate |
| near 50 percent. That is not enough supervision, and no amount of architecture fixes it. |
|
|
| It is also the most practically useful thing in this repo. Simulated syndromes are free. |
| Hardware shots are not, and the shot budget is what actually limits learned decoders on real |
| machines. |
|
|
| ## Scope |
|
|
| Distance 3 and 5, one code, one crosstalk model. The architectures are standard; the |
| measurement is the part that is mine. State of the art is AlphaQubit from Google DeepMind, |
| which used far more compute and real device data, and nothing here competes with it. |
| Regimes A and B are simulated, so leakage and calibration drift are absent. The distance-5 |
| models remain budget-limited. |
|
|
| ## License |
|
|
| MIT for code and weights. The device evaluation uses data released by Google Quantum AI under |
| CC-BY-4.0; attribution is on the dataset card. |
|
|