Zhejiang University · DCD Lab · Hangzhou

Yuhan Wang

Undergraduate Researcher · Multimodal Speech Intelligence

I study how machines can perceive, generate, and interact through sound. My recent work spans spoken-dialogue reward modeling, multimodal spatial audio, and full-duplex spoken interaction.

  • Spoken Dialogue
  • Spatial Audio
  • Full-Duplex Interaction
Illustrated avatar of Yuhan Wang reading a book
Yuhan Wang Computer Science · ZJU Hangzhou, China

Background

About

Building academically grounded audio systems from data and modeling to evaluation.

I am an undergraduate student in Computer Science and Technology at Zhejiang University, also pursuing the Advanced Honors Class of Engineering Education at Chu Kochen Honors College.

Since September 2024, I have worked as a research assistant in the DCD Lab under the supervision of Prof. Zhou Zhao. My current work connects spoken-language intelligence with generative audio: learning better reward signals for natural dialogue, building large-scale multimodal audio data, and developing controllable spatial and full-duplex audio systems.

Research output

Selected Publications

Peer-reviewed work in spoken dialogue, spatial audio, and singing voice synthesis, with links to official records and open research resources.

  1. EMNLP 2026 · Main Conference Second author · Survey

    Speaking While Listening: A Survey and Empirical Audit of Full-Duplex Spoken Dialogue Systems

    Jingyu Lu, Yuhan Wang, Jianming Luo, Yifu Chen, Tianle Liang, Shengpeng Ji, Ziyue Jiang, Xiaoda Yang, Yu Zhang, Xize Cheng, Chenyuhao Wen, Changhao Pan, Haoxiao Wang, Chen Ye, Jian Wu, Xiaoxi Jiang, Guanjun Jiang, Zhou Zhao†

    L0 to L3 architectural hierarchy showing where a full-duplex spoken dialogue system makes its listen, speak, wait, or dual decision

    A survey and empirical audit that separates where duplex decisions are made, which interaction is occurring, and how system behavior evolves moment by moment.

    I developed the L0–L3 architectural hierarchy, the T×I×R interaction ontology, and a five-state decision machine, and audited representative models, training datasets, and evaluation benchmarks.

  2. NeurIPS 2025 · Datasets & Benchmarks Dataset & benchmark

    MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations

    Wenxiang Guo*, Changhao Pan*, Zhiyuan Zhu*, Xintong Hu*, Yu Zhang*, Li Tang, Rui Yang, Han Wang, Zongbao Zhang, Yuhan Wang, Yixuan Chen, Hankun Xu, Ke Xu, Pengfei Fan, Zhetao Chen, Yanhao Yu, Qiange Huang, Fei Wu, Zhou Zhao†

    MRSAudio overview showing the MRSSpeech, MRSLife, MRSMusic, and MRSSing subsets with their multimodal spatial audio annotations

    A 484-hour multimodal spatial-audio dataset spanning speech, singing, music, and everyday scenes, with synchronized audio, video, trajectories, and refined annotations.

    I contributed to data collection and processing, audio-video and 3D-position synchronization, preprocessing, and speech-text alignment.

  3. Findings of ACL 2025 Singing voice synthesis

    TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

    Yu Zhang*, Wenxiang Guo*, Changhao Pan*, Dongyu Yao, Zhiyuan Zhu, Ziyue Jiang, Yuhan Wang, Tao Jin, Zhou Zhao†

    TCSinger 2 architecture with style transfer, blurred boundary content encoder, and flow-based custom Transformer modules

    A multilingual zero-shot singing voice synthesis model supporting style transfer and multi-level control from audio and natural-language prompts.

    I contributed to multilingual dataset expansion and annotation, Custom Audio Encoder validation, and contrastive-data preprocessing.

* Equal contribution. † Corresponding author. Publication links point to the official proceedings record whenever available.

Research agenda

Research Directions

One coherent direction: making audio agents more perceptive, generative, and natural in real-time interaction.

01

Spoken Dialogue Learning

Reward modeling and preference data for modality-aware, colloquial, multi-turn dialogue.

02

Spatial Audio Generation

Multimodal generation, refined annotation, and evaluation for multi-source spatial scenes.

03

Full-Duplex Interaction

Low-latency spoken systems that can listen, respond, and coordinate naturally in real time.

Ongoing research · 2026–Present Spoken interaction

Learning for Full-Duplex Spoken Dialogue

Studying evaluation and learning methods for spoken systems that must listen, respond, and manage turn-taking continuously rather than alternate between fixed user and assistant turns.

  • Turn-taking and interruption-aware evaluation
  • Long-context conversational learning
View the survey project
Zhejiang University Natural Science Cultivation Program · 2025–2026 One of 10 university-wide projects

Multimodal Multi-Source Spatial Audio Synthesis

Contributed to the ISDrama end-to-end spatial-audio architecture and the 484-hour MRSAudio dataset, spanning data collection, cross-modal synchronization, preprocessing, and refined annotation.

  • Audio-video and 3D-position synchronization
  • Speech-text alignment and dataset processing
View MRSAudio

Academic & industry

Experience

Research, industry, open-source, and academic experience in speech and machine learning.

May 2026 — Present

Industry research

Speech Algorithm Intern · Qwen Business Unit of Alibaba

Alibaba Group

Research and evaluation of end-to-end real-time spoken dialogue models, with emphasis on concurrent listening and speaking, interruption handling, turn transitions, and multi-turn interaction quality.

Jan 2026 — Apr 2026

Open source

Contributor · RLinf

Reinforcement Learning Infrastructure for Embodied and Agentic AI

Implemented a complete D4RL offline-IQL training and checkpoint-resume workflow in PyTorch, with dataset, Runner, Worker, and multi-task configuration support for MuJoCo, AntMaze, Kitchen, and Adroit tasks.

Sep 2024 — Present

Research

Research Assistant · DCD Lab

Zhejiang University · Supervisor: Prof. Zhou Zhao

Research on spatial audio generation, audio and speech processing, spoken-dialogue modeling, and full-duplex audio interaction.

2023 — 2027 (expected)

Education

Undergraduate Student · Computer Science and Technology

Zhejiang University

Minor · Advanced Honors Class of Engineering Education, Chu Kochen Honors College

Recognition

Selected Honors

A concise selection of scholarships and academic awards.

2024 & 2025

National Scholarship

Awarded in two consecutive academic years · Top 3%

2024–2025

Top 10 Campus Star

College of Computer Science and Technology · 10 / 1,305

2024–2025

Tencent Scholarship

One of five recipients in the college

2024

Zhejiang Provincial Higher Mathematics Competition · First Prize

Engineering category

2024

Zhejiang Provincial Physics Competition · Second Prize

Theoretical category