学术讲座公告---Advanced Methods in Regional Speech Enhancement and Far-to-Near-Field Transformation-武汉大学电子信息学院

信誉的新葡京官网

学术讲座

学术讲座公告---Advanced Methods in Regional Speech Enhancement and Far-to-Near-Field Transformation

2026-08-26
  • 阅读:

报告题目:Advanced Methods in Regional Speech Enhancement and Far-to-Near-Field Transformation

报告人: 于蒙 博士(腾讯混元大模型部门)

邀请人: 黄公平 教授

报告时间: 2026/8/27(周四)下午15:00

报告地点: 武汉大学电子信息学院 于刚·宋晓楼 A603会议室

报告简介:

Enhancing speech signals in complex acoustic environments remains a persistent challenge in audio processing. In recent work, the presenter introduces several innovative strategies to address core issues in this field. First, they propose a deep learning–based audio zooming technique that moves beyond conventional direction-dependent beamforming, instead enabling sound capture within a user-adjustable 3D region. This design allows for more precise and adaptable audio acquisition, supporting real-time use cases such as remote conferencing, education, and live streaming.

Building on this region-based capture paradigm, the work further aims to convert far-field audio into near-field quality. Using real-world acoustic data, the authors develop a novel framework that combines a Schrödinger Bridge–based diffusion model with generative adversarial networks. This hybrid approach achieves state-of-the-art performance in jointly suppressing noise and reverberation, leading to substantial improvements in speech quality. The proposed method sets a new benchmark for far-field to near-field enhancement in practical scenarios. Collectively, these advancements provide robust solutions to long-standing speech processing challenges, paving the way for high-fidelity audio experiences across a wide range of applications.

报告人简介:

Meng Yu has been a Principal Research Scientist at Tencent AI Lab, now at Tencent Hunyuan since 2016. From 2013 to 2016, he worked as a Staff Research Engineer at Audience, focusing on audio/speech enhancement for voice communication and improving speech recognition. Prior to that, he was a Software Engineer at Cisco from 2012 to 2013, specializing in speaker segmentation and recognition. He received B.S. in Mathematics from Peking University, Beijing, China in 2007, and a Ph.D. degree in Mathematics from University of California, Irvine, CA, USA in 2012. His research interests focus on audio and speech processing, with a particular emphasis on single and multi-channel far-field frontend speech enhancement applications.

欢迎感兴趣的老师和同学们积极参与!


学院地址:湖北省武汉市武昌区八一路299号(430072)

Address:No.299 Bayi Road,Wuhan,Hubei(P.R.C.:430072)

联系电话(Tel):(+86)27-68756275/68778537

传真(Fax):(+86)27-68778537

网址(Http):Http://eis.whu.edu.cn

联系邮箱(Email):[email protected]

武汉大学电子信息学院

官方微信公众号


©Copyright 2026 武汉大学电子信息学院  版权所有

新葡京在线网站-新葡京线上赌博平台 信誉的新葡京官网-新葡京在线直播投注 在线新葡京官网-新葡京靠谱的网上赌博平台 金好运娱乐城官网-金好运赌博平台 金好运老虎机官网-金好运娱乐城平台