Heterogeneous Computing for Real-Time Quantum State Classification
Under reviewA hardware-aware LSTM accelerator on an AMD Versal device for superconducting qubit readout. Independent input projections run on the AI Engine array while the loop-carried recurrent update stays in programmable logic. On a qubit with mid-readout state transitions, the design matches a full-resolution feed-forward network (96.59% vs. 96.37%) with 403× fewer parameters and about 103× fewer DSPs, and a 68.2% relative reduction in error over GMM, 43.6% over QubiCML, and 16.82% over MF-NN. Using 4 of 400 AIE tiles and 17 DSP58, it needs 2.2–8.2× fewer DSP58 than QubiCML and MF-NN; the AIE–PL split cuts tail latency by 21% and meets the sub-μs latency budget, using 31.1× fewer DSP58 than an all-PL implementation. A compact variant reaches 96.78% mean accuracy across eight qubits, the highest among all classifiers.