Developing advanced deep learning models for high-fidelity speech synthesis and voice generation. We focus on creating natural, expressive, and controllable human speech while preserving precise speaker identity and acoustic nuances.
Extending large language models beyond text to natively understand and generate speech. We explore unified architectures that seamlessly integrate audio and language modalities to build more intuitive and interactive AI systems.