y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto

#kv-cache-reduction News & Analysis

2 articles tagged with #kv-cache-reduction. AI-curated summaries with sentiment analysis and key takeaways from 50+ sources.

2 articles
AIBullisharXiv – CS AI · Jun 97/10
🧠

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech

Researchers introduce TLDR, a patch-based autoregressive framework that compresses audio tokens to accelerate text-to-speech synthesis. The method achieves 1.8x inference speedup and reduces KV-cache memory by 75% without replacing existing model modules, addressing a key efficiency bottleneck in codec-based speech language models.

AIBullisharXiv – CS AI · Jun 47/10
🧠

Do Transformers Need Three Projections? Systematic Study of QKV Variants

Researchers systematically evaluate whether transformer models require three separate QKV projections, discovering that shared projection variants perform comparably while reducing computational overhead. The Q-K=V configuration achieves 50% KV cache reduction with minimal performance loss and combines effectively with existing optimization techniques like MQA to enable practical on-device deployment.

🏢 Perplexity