Vu Le

PhD Student · UMass Amherst · Berkeley Lab

I'm a third-year Computer Science PhD student at the University of Massachusetts Amherst (UMass), working with Prof. VP Nguyen and Prof. Deepak Ganesan. I am also an affiliated PhD student at Berkeley Lab, where I work on computer architecture and heterogeneous computing for real-time quantum systems, including machine learning-based quantum state classification under a sub μs latency budget, hosted by Dr. Yilun Xu.

I am broadly interested in HPC, computer architecture, and heterogeneous computing, with a focus on hardware-software co-design for real-time AI inference and scientific computing. My research explores how we can exploit key features of AI Engines, FPGAs, CPUs, GPUs and emerging hardware to efficiently execute compute-intensive & memory bounded workloads under strict latency and resource constraints, with applications in quantum computing and other performance-critical systems.

I am seeking research or industry internships in the US for Spring 2027/Fall 2027, particularly in AI acceleration, edge computing, and heterogeneous computing. Feel free to reach out!

Get in touch

Personal: vule20.cs AT gmail [DOT] com
UMass: vdle AT cs.umass [DOT] edu

CV (PDF)

Active Research

AIE–PL LSTM pipeline for real-time superconducting qubit readout

Heterogeneous Computing for Real-Time Quantum State Classification

Under review

A hardware-aware LSTM accelerator on an AMD Versal device for superconducting qubit readout. Independent input projections run on the AI Engine array while the loop-carried recurrent update stays in programmable logic. On a qubit with mid-readout state transitions, the design matches a full-resolution feed-forward network (96.59% vs. 96.37%) with 403× fewer parameters and about 103× fewer DSPs, and a 68.2% relative reduction in error over GMM, 43.6% over QubiCML, and 16.82% over MF-NN. Using 4 of 400 AIE tiles and 17 DSP58, it needs 2.2–8.2× fewer DSP58 than QubiCML and MF-NN; the AIE–PL split cuts tail latency by 21% and meets the sub-μs latency budget, using 31.1× fewer DSP58 than an all-PL implementation. A compact variant reaches 96.78% mean accuracy across eight qubits, the highest among all classifiers.

Hard real-time AI inference on AMD AI Engines

Hard-Real-Time AI Inference on AI Engines

In progress

A study of latency-bounded deep learning inference on spatial AI accelerators like AMD AI Engines, characterizing and modeling timing behavior to support deployments with strict real-time guarantees.

Selected Research

My research focuses on heterogeneous computing for real-time AI inference — combining AI engines, FPGAs, and accelerators to meet hard latency constraints. I'm broadly interested in computer architecture, quantum computing systems, deep learning, and scalable networked systems. Some papers are

Detection and Tracking of Drone Swarms using LiDAR 2025

Detection and Tracking of Drone Swarms using LiDAR

Tasnim Azad Abir, Vu Le, Endrowednes Kuantama, Pranjol Sen Gupta, Austin Copley, Judith Dawes, Mohammad Islam, Richard Han, Phuc Nguyen

ACM MobiSys 2025 · A* conference

LiSWARM is a low-cost LiDAR system for accurate 3D tracking and recognition of drones in large swarms. Using point cloud processing, clustering, and neural networks, it achieves up to 98% accuracy and scales to 15,000 drones—enabling applications in airspace security, drone shows, and sensitive area monitoring.

MagicStream immersive telepresence 2024

MagicStream: Bandwidth-conserving Immersive Telepresence via Semantic Communication

Ruizhi Cheng, Nan Wu, Vu Le, Eugene Chai, Matteo Varvello, Bo Han

ACM SenSys 2024 · A* conference

MagicStream, a first-of-its-kind semantic-driven immersive telepresence system that effectively extracts and delivers compact semantic details of captured 3D representation of users, instead of traditional bit-by-bit communication of raw content.

Fast and Interpretable Face Identification using Vision Transformers 2024

Fast and Interpretable Face Identification for Out-Of-Distribution Data Using Vision Transformers

Hai Phan, Cindy Le, Vu Le, Yihui He, Anh Totti Nguyen

CVF/WACV 2024 · A conference

Using vision transformers for out-of-distribution data face identification, runs twice faster while achieving comparable performance with the state of the art DeepFace-EMD model.

News

04/2026

Presented hardware-software co-design with spatial accelerators and FPGA for real-time qubit readout to LBNL and AMD in Berkeley, CA.

06/2025

Presented “Opportunities in Computer Systems Research for Quantum Computing” at ACM QSys 2025 (in conjunction with ACM MobiSys 2025).

03/2025

One paper accepted at ACM MobiSys 2025.

12/2024

I officially become a research affiliate with Berkeley Lab.

09/2024

My new academic website with the vule.us domain is live now.

09/2024

One paper accepted at ACM SenSys 2024.

09/2024

I joined University of Massachusetts Amherst, USA as a PhD student.

04/2024

I received some CS PhD offers in the US.

10/2023

One paper accepted at IEEE/CVF WACV.

Miscellanea

Apart from research, I'm an experienced software and DevOps engineer who enjoys building scalable backend systems. Outside of work, I'm an avid adventurer — I love road trips and have driven across the US to explore national parks and trails firsthand. Most of my hikes and adventures (Grand Canyon, Zion, Death Valley, Horseshoe Bend, Joshua Tree, the White Mountains in New Hampshire) were made possible by hitting the road. I also really enjoy lifting at the gym, jogging, brewing coffee, planting flowers, and skiing. I shoot with a Sony A6400 and a collection of Sigma and Sony lenses. Check out my photo gallery.