Zhejiang University · DCD Lab · Hangzhou
Yuhan Wang
Undergraduate Researcher · Multimodal Speech Intelligence
I study how machines can perceive, generate, and interact through sound. My recent work spans spoken-dialogue reward modeling, multimodal spatial audio, and full-duplex spoken interaction.
Background
About
Building academically grounded audio systems from data and modeling to evaluation.
I am an undergraduate student in Computer Science and Technology at Zhejiang University, also pursuing the Advanced Honors Class of Engineering Education at Chu Kochen Honors College.
Since September 2024, I have worked as a research assistant in the DCD Lab under the supervision of Prof. Zhou Zhao. My current work connects spoken-language intelligence with generative audio: learning better reward signals for natural dialogue, building large-scale multimodal audio data, and developing controllable spatial and full-duplex audio systems.
Research output
Selected Publications
Peer-reviewed work in spoken dialogue, spatial audio, and singing voice synthesis, with links to official records and open research resources.
-
-
Speaking While Listening: A Survey and Empirical Audit of Full-Duplex Spoken Dialogue Systems
A survey and empirical audit that separates where duplex decisions are made, which interaction is occurring, and how system behavior evolves moment by moment.
I developed the L0–L3 architectural hierarchy, the T×I×R interaction ontology, and a five-state decision machine, and audited representative models, training datasets, and evaluation benchmarks.
-
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations
A 484-hour multimodal spatial-audio dataset spanning speech, singing, music, and everyday scenes, with synchronized audio, video, trajectories, and refined annotations.
I contributed to data collection and processing, audio-video and 3D-position synchronization, preprocessing, and speech-text alignment.
-
TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
Research agenda
Research Directions
One coherent direction: making audio agents more perceptive, generative, and natural in real-time interaction.
Spoken Dialogue Learning
Reward modeling and preference data for modality-aware, colloquial, multi-turn dialogue.
Spatial Audio Generation
Multimodal generation, refined annotation, and evaluation for multi-source spatial scenes.
Full-Duplex Interaction
Low-latency spoken systems that can listen, respond, and coordinate naturally in real time.
Learning for Full-Duplex Spoken Dialogue
Studying evaluation and learning methods for spoken systems that must listen, respond, and manage turn-taking continuously rather than alternate between fixed user and assistant turns.
- Turn-taking and interruption-aware evaluation
- Long-context conversational learning
Multimodal Multi-Source Spatial Audio Synthesis
Contributed to the ISDrama end-to-end spatial-audio architecture and the 484-hour MRSAudio dataset, spanning data collection, cross-modal synchronization, preprocessing, and refined annotation.
- Audio-video and 3D-position synchronization
- Speech-text alignment and dataset processing
Academic & industry
Experience
Research, industry, open-source, and academic experience in speech and machine learning.
Industry research
Speech Algorithm Intern · Qwen Business Unit of Alibaba
Alibaba Group
Research and evaluation of end-to-end real-time spoken dialogue models, with emphasis on concurrent listening and speaking, interruption handling, turn transitions, and multi-turn interaction quality.
Open source
Contributor · RLinf
Reinforcement Learning Infrastructure for Embodied and Agentic AI
Implemented a complete D4RL offline-IQL training and checkpoint-resume workflow in PyTorch, with dataset, Runner, Worker, and multi-task configuration support for MuJoCo, AntMaze, Kitchen, and Adroit tasks.
Research
Research Assistant · DCD Lab
Zhejiang University · Supervisor: Prof. Zhou Zhao
Research on spatial audio generation, audio and speech processing, spoken-dialogue modeling, and full-duplex audio interaction.
Education
Undergraduate Student · Computer Science and Technology
Zhejiang University
Minor · Advanced Honors Class of Engineering Education, Chu Kochen Honors College
Recognition
Selected Honors
A concise selection of scholarships and academic awards.
SenseTime Scholarship
National cohort · 30 recipients
National Scholarship
Awarded in two consecutive academic years · Top 3%
Top 10 Campus Star
College of Computer Science and Technology · 10 / 1,305
Tencent Scholarship
One of five recipients in the college
Zhejiang Provincial Higher Mathematics Competition · First Prize
Engineering category
Zhejiang Provincial Physics Competition · Second Prize
Theoretical category