TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech
Researchers introduce TLDR, a patch-based autoregressive framework that compresses audio tokens to accelerate text-to-speech synthesis. The method achieves 1.8x inference speedup and reduces KV-cache memory by 75% without replacing existing model modules, addressing a key efficiency bottleneck in codec-based speech language models.