ByteDance introduces SwanTale audio generation model
ByteDance researchers have introduced SwanTale, a unified audio model that simplifies multi-speaker speech and sound generation for both natural language instructions and zero-shot cloning.

ByteDance researchers have introduced SwanTale, an expressive multi-speaker speech and audio generation model designed to handle both instruct and zero-shot tasks. The system addresses the complex demands of modern media production, such as animation dubbing, audio dramas, movies, advertising, games, podcasts, and short-video production. In these creative scenarios, developers often need to design unique voices without relying on existing reference recordings, control speaker styles using natural language, and support acoustic scenes with realistic environments and audio effects. By supporting both instruct and zero-shot tasks, SwanTale allows creators to generate high-quality audio and later reuse the designed voices. In instruct tasks, the model relies on natural language captions detailing speaker styles, fine-grained content, and environmental acoustics. In zero-shot tasks, it clones voices using a brief reference recording alongside the target content.
To achieve high-quality generation across multiple audio modalities, the developers built SwanTale using a combination of novel data and architectural techniques. On the data side, they created SwanData-Caption, a pipeline that cleans raw speech and audio data, injects targeted synthetic coverage, and annotates diverse and accurate multi-level captions. The model itself utilizes SwanVAE to manage multi-audio-modality outputs. It also incorporates reward-conditioned quality control, Engram conditioning, and a Unified Mixture of Experts (MoE) architecture to handle diverse tasks and modalities simultaneously.
The training process for SwanTale leverages curriculum learning and GRPO post-training, allowing the model to progressively learn and strengthen its capabilities. According to the research paper, this approach enables SwanTale to lead on multiple key zero-shot and instruct metrics. The model achieved the highest expressiveness scores across both task categories and successfully demonstrated complex instruct-based generation involving multiple speakers and environmental audio effects.
This is our own summary of reporting by HF Papers



