About VibeVoice
VibeVoice is an open-source community dedicated to advancing speech synthesis technology. Our goal is to make high-quality, long-form, multi-speaker speech generation accessible to everyone.
Open Source
We believe in the power of open source and are committed to building a transparent and collaborative community. MIT/Apache license for everyone.
Innovation
Constantly exploring the boundaries of LLM and Diffusion models in the speech domain, pushing speech synthesis forward.
Community Driven
Listening to users and growing together with developers to build better voice generation tools.
About the Project
VibeVoice is developed by Microsoft Research Asia. It is a speech generation model based on the next-token diffusion mechanism, capable of generating up to 90 minutes of high-quality audio with up to 4 speakers in natural conversation.
The project's core innovation lies in combining the context understanding capabilities of Large Language Models (LLM) with the high-fidelity audio generation capabilities of Diffusion models, achieving unprecedented speech naturalness and expressiveness.
Our goal is to make content creation for podcasts, audiobooks, video dubbing, and more simpler and more efficient, enabling everyone to easily create professional-grade voice content.