y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto

#model-compression News & Analysis

194 articles tagged with #model-compression. AI-curated summaries with sentiment analysis and key takeaways from 50+ sources.

194 articles
AIBullisharXiv – CS AI · Jun 196/10
🧠

SoftSkill: Behavioral Compression for Contextual Adaptation

SoftSkill introduces a method to compress natural-language AI agent skills into compact continuous context objects that improve task performance without retraining frozen language models. By replacing lengthy Markdown skill files with 32-token soft prefixes, the approach demonstrates significant accuracy gains across multiple benchmarks while reducing computational overhead.

AIBullisharXiv – CS AI · Jun 126/10
🧠

Reducing the Complexity of Deep Learning Models for EEG Analysis on Wearable Devices

Researchers demonstrate that deep learning models for EEG analysis can be significantly compressed through parameter quantization and electrode reduction techniques, enabling deployment on resource-constrained wearable devices without substantial accuracy loss. This addresses a critical bottleneck in portable healthcare technology where computational demands of DNNs far exceed device capabilities.

AINeutralarXiv – CS AI · Jun 116/10
🧠

SPEAR: A System for Post-Quantization Error-Adaptive Recovery Enabling Efficient Low-Bit LLM Serving

SPEAR is a new system that improves efficiency of quantized large language models by using adaptive error correction tailored to individual tokens, rather than static corrections applied uniformly. The technique recovers 56-75% of the performance gap between 4-bit and full-precision models while adding minimal memory overhead, advancing practical LLM deployment at scale.

🏢 Perplexity
AINeutralarXiv – CS AI · Jun 106/10
🧠

What Fits (Into Few Tokens) Doesn't Overfit: Compression and Generalization in ML Research Agents

Researchers demonstrate that successful machine learning strategies remain highly compressible and generalizable even when trained on held-out benchmarks, suggesting overfitting in benchmark-driven ML is rare because effective strategies occupy a low-complexity region of strategy space. Using LLM-driven research agents, they show that short prompts and minimal feedback suffice to reproduce high-performance models across diverse domains.

AIBullisharXiv – CS AI · Jun 106/10
🧠

ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models

ReasonAlloc is a training-free framework that optimizes key-value cache memory allocation during LLM inference for reasoning tasks by using hierarchical, non-uniform budget distribution across layers and attention heads. The method significantly reduces memory bottlenecks in chain-of-thought reasoning while maintaining performance, outperforming existing compression approaches on mathematical reasoning benchmarks.

🧠 Llama
AINeutralarXiv – CS AI · Jun 106/10
🧠

Interactions Between Crosscoder Features: A Compact Proofs Perspective

Researchers introduce a framework using compact proofs to measure feature interactions in crosscoders and Sparse Autoencoders, revealing that interactions between learned features cause reconstruction errors. The work demonstrates practical applications including computationally sparse models that maintain 60% performance with minimal features and detection of sleeper agent behavior in AI systems.

AINeutralarXiv – CS AI · Jun 96/10
🧠

Structured Neuron Pruning in Deep Neural Networks Using Multi-Armed Bandits

Researchers present a novel structured pruning framework that uses multi-armed bandit algorithms to remove redundant neurons from deep neural networks. The approach treats each neuron as a bandit arm, testing its importance through temporary masking and loss measurement, then applies various MAB policies (UCB1, Thompson Sampling, etc.) to identify which neurons to prune. Experiments across tabular and deep learning tasks show MAB-based pruning significantly outperforms traditional magnitude-based and greedy pruning methods.

AINeutralarXiv – CS AI · Jun 96/10
🧠

Seq103: A Unified Neuroevolution Framework for Compact Sequence Architecture Discovery

Seq103 introduces a unified neuroevolution framework that automatically discovers compact neural network architectures for sequence tasks, achieving 81-87% of baseline accuracy while using 11-3,200x fewer parameters. The framework applies the same evolutionary search pipeline to both recurrent and feedforward sequence classification, offering significant efficiency gains for resource-constrained deployments.

AINeutralarXiv – CS AI · Jun 96/10
🧠

Understanding Quantization-Aware Training: Gradients at Quantized Weights Bias to the Low-Loss Basin

Researchers propose a geometric framework explaining why post-training quantization (PTQ) fails at aggressive bitwidths while quantization-aware training (QAT) succeeds in recovery. The study reveals that gradients in QAT acquire an inward bias toward low-loss regions, enabling quantized neural networks to maintain accuracy where simpler PTQ methods collapse.

AINeutralarXiv – CS AI · Jun 96/10
🧠

Hyperflux: Pruning Reveals Importance

Researchers introduce Hyperflux, a novel L0 pruning method that models neural network pruning as a dynamically evolving system driven by flux and pressure mechanisms. The approach provides interpretability at multiple scales while achieving competitive sparsity results on standard vision benchmarks, advancing understanding of how neural networks can be efficiently compressed.

AIBullisharXiv – CS AI · Jun 96/10
🧠

Learning Quantized Continuous Controllers for Integer Hardware

Researchers demonstrate quantization-aware training techniques that compress reinforcement learning policies to 2-3 bits per weight while maintaining performance comparable to full-precision models, enabling efficient deployment on resource-constrained FPGA hardware with microsecond-level inference latency.

AIBullisharXiv – CS AI · Jun 96/10
🧠

Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs

Researchers introduce Ghosted Layers, a training-free method to recover performance degradation in layer-pruned large language models by solving an activation alignment problem through optimal linear operators. The technique uses a small calibration set to reconstruct hidden state mismatches introduced by pruning, maintaining efficiency gains while improving accuracy and perplexity across multiple LLM architectures.

🏢 Perplexity
AINeutralarXiv – CS AI · Jun 56/10
🧠

LoRi: Low-Rank Distillation for Implicit Reasoning

Researchers propose LoRi, a low-rank distillation framework that improves implicit chain-of-thought reasoning in large language models by aligning teacher-student model trajectories in a shared low-rank tensor subspace. The method addresses the performance gap between implicit and explicit reasoning approaches, showing consistent improvements across LLaMA and Qwen model families on mathematical benchmarks.

AINeutralarXiv – CS AI · Jun 56/10
🧠

Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads

Researchers have identified a structural property in Multimodal Large Language Models called functional sparsity, discovering specialized attention heads (CoRe heads) that efficiently extract relevant visual information from complex contexts. This mechanistic insight demonstrates that only the top 5% of these heads are critical for multimodal reasoning, suggesting significant potential for model optimization and inference acceleration without performance loss.

AINeutralarXiv – CS AI · Jun 56/10
🧠

Surrogate Neural Architecture Codesign Package (SNAC-Pack)

SNAC-Pack is an open-source AutoML framework that automates neural architecture design for FPGA deployment by combining hardware-aware search with quantization and pruning. The tool reduces design cycles from months to hours while matching or exceeding baseline performance on tasks like jet classification and quantum computing applications.

AINeutralarXiv – CS AI · Jun 45/10
🧠

Gravity-Aware Hierarchical Routing for Lightweight SensorLLM on Human Activity Recognition

Researchers propose a gravity-aware hierarchical routing method to improve human activity recognition in compressed language models used with wearable sensors. The lightweight adaptation addresses a specific failure mode where static activities like standing and sitting are poorly recognized when using compact models like TinyLlama, while maintaining strong performance on dynamic activities.

AINeutralarXiv – CS AI · Jun 46/10
🧠

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats

Researchers introduce dMX, a differentiable mixed-precision quantization framework that enables dynamic floating-point bit-width assignment across different layers of large language models. The method uses continuous optimization with temperature-based annealing to efficiently compress models while maintaining accuracy, demonstrating improvements over existing quantization heuristics across multiple LLM families.

🏢 Perplexity🧠 Llama
AIBullisharXiv – CS AI · Jun 46/10
🧠

MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models

Researchers introduce MorphoQuant, a post-training quantization framework designed to compress omni-modal large language models to 4-bit precision while preserving cross-modal performance. The method addresses distribution heterogeneity across different data modalities through bias compensation and quantization grid optimization, achieving results that rival higher-precision baselines.

AINeutralarXiv – CS AI · Jun 46/10
🧠

Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision Transformers

Researchers introduce MaskAQ, a novel data-free quantization technique for Vision Transformers that identifies and aligns informative image regions to improve model compression without requiring access to real training data. The approach addresses distribution mismatches in synthetic data generation, enabling more efficient deployment of ViT models while maintaining security and privacy.

AIBullisharXiv – CS AI · Jun 26/10
🧠

Logit Distillation on Manifolds: Mapping by Learning

Researchers introduce a layer-wise projection mapping technique for knowledge distillation that enables efficient model compression, reducing trainable parameters to under 1% of the teacher model while maintaining performance improvements. Combined with LoRA injection, this approach significantly outperforms traditional distillation methods in word error rate metrics and enables rapid parallel training without the computational overhead of mixture-of-experts models.

AINeutralarXiv – CS AI · Jun 26/10
🧠

DASH: Dual-Branch Score Distillation for Guidance-Calibrated Compact Diffusion Models

DASH introduces a dual-branch distillation framework for compressing class-conditional diffusion models while preserving classifier-free guidance effectiveness. By independently supervising both conditional and unconditional score branches, the method achieves 5.9x model compression with minimal quality degradation, addressing a critical limitation in existing distillation approaches where guidance mechanisms collapse during compression.

AINeutralarXiv – CS AI · Jun 26/10
🧠

What Makes a Strong Model? A Unified Spectral Analysis of Knowledge Transfer over High-dimensional Linear Regression

Researchers present a unified theoretical framework analyzing knowledge transfer (KT) in machine learning through spectral analysis of SGD dynamics. The study reveals two distinct mechanisms—Spectral Horizon Expansion in knowledge distillation and Spectral Denoising in weak-to-strong generalization—explaining how knowledge transfer efficiency is governed by implicit regularization and heterogeneous spectral learning speeds.

AINeutralarXiv – CS AI · Jun 26/10
🧠

You Can Learn Tokenization End-to-End with Reinforcement Learning

Researchers propose learning tokenization boundaries in large language models using reinforcement learning and score function estimates instead of hardcoded compression. This approach directly optimizes discrete token boundaries, outperforming prior straight-through estimation methods at the 100 million parameter scale.

AINeutralarXiv – CS AI · Jun 25/10
🧠

Balancing Knowledge Distillation for Imbalance Learning with Bilevel Optimization

Researchers introduce BiKD, a bilevel optimization framework that dynamically adjusts the balance between hard and soft losses in knowledge distillation for imbalanced datasets. The method uses a weight generation network guided by a balanced validation set to assign per-sample adaptive weights, significantly improving performance on long-tailed datasets like CIFAR-10/100 compared to existing approaches.

← PrevPage 5 of 8Next →