About VibeVoice

VibeVoice is an open-source community dedicated to advancing speech synthesis technology. Our goal is to make high-quality, long-form, multi-speaker speech generation accessible to everyone.

πŸ”“

Open Source

We believe in the power of open source and are committed to building a transparent and collaborative community. MIT/Apache license for everyone.

πŸš€

Innovation

Constantly exploring the boundaries of LLM and Diffusion models in the speech domain, pushing speech synthesis forward.

πŸ‘₯

Community Driven

Listening to users and growing together with developers to build better voice generation tools.

About the Project

VibeVoice is developed by Microsoft Research Asia. It is a speech generation model based on the next-token diffusion mechanism, capable of generating up to 90 minutes of high-quality audio with up to 4 speakers in natural conversation.

The project's core innovation lies in combining the context understanding capabilities of Large Language Models (LLM) with the high-fidelity audio generation capabilities of Diffusion models, achieving unprecedented speech naturalness and expressiveness.

Our goal is to make content creation for podcasts, audiobooks, video dubbing, and more simpler and more efficient, enabling everyone to easily create professional-grade voice content.