Date: Friday, Sep 4, 2026
Time: 2:00 -4:00 PM
Venue: IB 3106
Zoom: 789 843 9496

Speaker

Dr. Zhe Li
The University of Hong Kong, Postdoctoral Fellow
Abstract
Pre-trained speech foundation models provide broadly transferable representations, but downstream adaptation raises two complementary questions: what should be learned during post-training, and where should task-specific capacity be introduced? This talk presents a unified view of efficient and reliable post-training. The first part examines learning objectives for speech LLMs. Supervised fine-tuning establishes multimodal task alignment, while reinforcement learning with verifiable rewards promotes accurate, structured, and evidence-grounded outputs. Component-wise reward normalization and grounding-aware advantage gating improve credit assignment across correctness, formatting, and grounding signals. The second part examines efficient adaptation across three spaces: dynamic prompts for input conditioning, mixture-of-experts adapters for intermediate representations, and spectral low-rank adaptation for pre-trained weights. These methods show how the placement and structure of a small number of trainable parameters affect efficiency, task specialization, and generalization. These studies frame post-training as the joint design of what models should learn and which model components should be adapted.
Bio
Zhe Li is a Postdoctoral Fellow at The University of Hong Kong and will join Stanford University as a Postdoctoral Fellow in December 2026. His research interests include speech LLMs, robust speaker representation learning, and multimodal artificial intelligence for healthcare applications. He received his PhD from The Hong Kong Polytechnic University in 2025. He was a research intern at Microsoft Research Asia in 2025 and a visiting PhD researcher at Stanford University in 2024. He has published more than 60 papers in leading speech journals and conferences, including IEEE TASLP, ICASSP, and INTERSPEECH. He holds three granted patents. He delivered tutorials on speech LLMs at ICME 2026 and INTERSPEECH 2026 and co-organized a special session at ACM MMAsia 2026. He received the 2020 Outstanding Scientific and Technological Achievement Award from the Chinese Association for Artificial Intelligence. His co-authored work received the Best Student Paper Runner-Up Award at PRICAI 2024.
Speaker

Dr. Zezhong Jin
The Hong Kong Polytechnic University
Abstract
Knowledge distillation provides an effective way to transfer capabilities from large teacher models to compact students. In this talk, I will present three studies on knowledge distillation for speaker verification, focusing on label-level supervision, intermediate feature transfer, and attention distillation. These works examine what knowledge should be transferred and how teacher–student mismatch influences the effectiveness of distillation. I will then briefly connect these ideas to recent LLM distillation methods. The discussion will focus on two representative paradigms: offline distillation, where students learn from teacher-generated responses and reasoning traces, and on-policy distillation, where teachers provide feedback on student-generated trajectories. By connecting these settings, the talk develops a unified view of distillation as the transfer of increasingly rich knowledge, from predictions and representations to reasoning processes and interactive behaviors.
Bio
Zezhong Jin received his B.Eng. degree in Electronic Information Engineering from Hebei University in 2021, and his M.Sc. and Ph.D. degrees in Electronic and Information Engineering from The Hong Kong Polytechnic University in 2023 and 2026, respectively. His research interests include speaker recognition, speaker representation learning, speech large language models, and trustworthy speech intelligence. He has gained extensive research and industrial experience through internships at Microsoft Research Asia (MSRA), ByteDance under the Jindouyun Research Internship Program, Tencent, and Zoom.
He has published more than 15 papers in leading journals and conferences, including IEEE/ACM TASLP, EMNLP, Neurocomputing, ICASSP, and INTERSPEECH. His research spans speaker modeling, knowledge distillation, self-supervised speech representation learning, speech generation, and speech and multimodal intelligence. He has also actively contributed to the speech research community, serving as a reviewer for major international conferences including ICASSP and INTERSPEECH. In addition, he has delivered tutorials at international conferences, including ICME, on topics related to speaker representation learning and speech intelligence.