creators
Song Han: efficient AI at MIT and NVIDIA
Song Han is a tenured associate professor at MIT EECS and a research director at NVIDIA, known for Deep Compression, TinyML and LLM quantization.
Song Han is an associate professor with tenure at MIT EECS, where he leads the HAN Lab. NVIDIA's own research site describes him as a research director there, leading its Efficient AI team and co-leading the NVIDIA Singapore Lab. He received his S.B. in Microelectronics Engineering from Tsinghua University, and his SM and PhD in Electrical Engineering from Stanford, where he was advised by Bill Dally. His work is on making models cheap enough to run: Deep Compression, the pruning-and-quantization method, and the Efficient Inference Engine, which his MIT page says first brought weight sparsity to modern AI chips and which ranks among the five most cited papers in the fifty-year history of ISCA. The same line of work produced hardware-aware neural architecture search (Once-for-All), TinyML for constrained devices, and more recently LLM quantization and long-context inference - SmoothQuant, AWQ and StreamingLLM, all three taken up in NVIDIA's TensorRT-LLM. He teaches the open lecture series EfficientML.ai. He has best paper awards from ICLR 2016, FPGA 2017, MLSys 2024 (for AWQ) and MLSys 2026, an NSF CAREER Award, a 2023 Sloan Research Fellowship and MIT Technology Review's "35 Innovators Under 35". He is chairing MLSys 2027.
- Specialization
- efficient AI computing, model compression and quantization, TinyML, AI hardware
- Country
- United States