arXiv cs.CLPaper
Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis
The alignment-free approach and scale are solid improvements over Audiobox. Removing forced alignment reduces the error cascade in speech synthesis. This matters if you're building voice products, less if you're consuming APIs. The 3B parameter model trained on 480k hours signals meaningful engineering effort but doesn't change competitive dynamics unless it ships and performs at scale.