Yuxuan Wang

I’ve led foundational AI research, built multidisciplinary organizations, and turned research advances into widely adopted products.

I am currently Head of Speech & Realtime Agents at ByteDance Seed. Our research and products span speech, conversational AI, real-time agents, native multimodal interaction, and generative media.

I built the organizations behind these efforts from the ground up, bringing model research, data, infrastructure, and product together to turn early research bets into models deployed at scale.

Across ByteDance’s own products, our models serve hundreds of millions of daily active users and handle billions of API calls per day. They power AI assistants, content platforms, and creative and productivity tools, with further deployments in devices and enterprise applications. They are shaping how people interact, converse, and create with AI.

Earlier, I co-led Tacotron at Google and, before that, pioneered deep learning for speech separation. These contributions drove two paradigm shifts in speech AI.

I received my Ph.D. in Computer Science from The Ohio State University in 2015, where I was a Presidential Fellow.

Selected work

Native multimodal interaction

Our research spans multimodal interaction, full-duplex conversation, conversational AI, and real-time agents, with related work on memory, personalization and efficient inference. All models below are generally available in products.

  • SeedRealtime (2026) — An omni full-duplex interaction model with proactive engagement, natural conversational timing, and tool use.

  • Seeduplex (2026) — Native full-duplex speech-to-speech interaction, with listening, speaking, turn-taking, and interruptions handled jointly within the model.

  • Doubao Realtime Voice (2025) — Unifies speech understanding and generation in one native model, with reasoning, memory, emotional understanding, and expressive conversation.

  • Seed LiveInterpret 2.0 (2025) — End-to-end simultaneous interpretation that listens and translates concurrently while preserving the speaker’s voice.

Generative media

Our models generate speech, sound, and music for media creation, guided by multimodal context.

  • Seed Audio (2026) — A unified model for controllable audio creation, combining dialogue, sound effects, and ambience in coherent scenes. Our technology also powers audio generation in Seedance (2026).

  • Seed Music (2024) — A unified model for music generation and editing, with control over style, lyrics, melody, and vocal performance.

Speech & audio foundations

Our foundation models bring context, reasoning, and expressive control to audio perception and speech generation.

  • Audio perception (Seed2.0, 2026) — Native audio modeling that unifies a broad range of speech and audio understanding tasks, with joint reasoning across audio, text, and images.

  • Seed-ASR (2024) — LLM-based speech recognition with contextual understanding, accurate recognition of long-tail vocabulary, broad language and accent coverage, and robustness to noise. The technology proved so effective that we built a standalone product around it: Doubao IME.

  • Seed-TTS (2024) — Introduced in-context learning and reinforcement learning for expressive, controllable speech generation, alongside diffusion-based speech editing.

  • Tacotron (2016–2018) — Established the end-to-end neural paradigm for speech synthesis—the foundation of modern speech generation.

  • Deep learning for speech (2010–2014) — Established supervised deep learning as a new paradigm for speech separation, with problem formulations and training methods that underpin the field today.

    2019 IEEE Signal Processing Society Best Paper Award