AIBullisharXiv – CS AI · Jun 197/10
🧠Researchers introduce BA-solver, a lightweight acceleration method for Flow Matching generative models that achieves quality comparable to 100+ neural function evaluations using only 10 evaluations. The approach combines a frozen backbone model with a minimal SideNet (1-2% additional parameters) to approximate velocities bidirectionally, enabling faster image generation while maintaining compatibility with existing pipelines.
AIBullisharXiv – CS AI · May 277/10
🧠Researchers introduce FAV, a novel framework for aligning few-step generative models that requires only sample access to generators and reference distributions. The method uses Stein Variational Gradient Descent to cast alignment as sampling from reward-tilted distributions, demonstrating superior performance across robotic manipulation tasks and scaling to high-resolution image synthesis.
AIBullisharXiv – CS AI · Jun 236/10
🧠Researchers present Gazer, a training-free framework that uses multimodal large language models to identify and correct semantic errors in autoregressive visual models during image and video generation. The approach operates through diagnostic and correction stages that analyze intermediate generation states and adjust trajectories without requiring additional model training.
AINeutralarXiv – CS AI · Jun 236/10
🧠Researchers introduce Semantic Browsing, a method that improves diversity in AI-generated images by controlling variation at the text level rather than through random pixel-level changes. Using Vision Language Models and structured prompting, the technique enables users to explore meaningful, interpretable variations of generated images organized along semantic axes.
AINeutralarXiv – CS AI · Jun 106/10
🧠Researchers have developed a deep learning system that synthesizes intermediate CT slices to reduce through-plane anisotropy in head CT imaging, effectively halving spacing while simultaneously denoising outputs. The system outperforms classical interpolation and existing video frame interpolation methods, with MS-SSIM+L1 loss providing optimal performance across structural measures.
AINeutralarXiv – CS AI · Jun 106/10
🧠Researchers propose CCE-Diffusion, a framework that improves text-driven image generation by customizing concept embeddings to better align foreground objects with background synthesis. The method reduces visual artifacts in AI-generated product images, offering merchants a cost-effective tool for creating high-quality display content.
AINeutralarXiv – CS AI · Jun 95/10
🧠Researchers introduce BLM-SGAN, a novel text-to-image generation model that combines bidirectional language modeling with GANs to improve image synthesis from text descriptions. The model achieves state-of-the-art performance metrics, outperforming existing approaches by better capturing contextual dependencies and reducing training limitations.
AINeutralarXiv – CS AI · Jun 86/10
🧠Researchers introduce DIRECT, a novel framework for 3D-aware object insertion that combines interactive pose control with diffusion-based image synthesis. By decomposing insertion conditions into appearance, geometry, and context guidance through separate pathways, the method achieves superior control over object positioning and visual quality compared to existing 2D inpainting approaches.
AIBullisharXiv – CS AI · Jun 26/10
🧠Researchers propose VRPO, a reinforcement learning-based optimization method that improves training efficiency in diffusion transformers by dynamically aligning generative and discriminative representations. The approach replaces static alignment losses with adaptive reward-based optimization, achieving up to 1.8 FID improvement and 2.3x faster training compared to existing methods.
AINeutralarXiv – CS AI · May 296/10
🧠Researchers propose a unified deep learning framework that synthesizes virtual monochromatic 50 keV CT images from standard single-energy CT scans by conditioning on contrast phase information. This approach addresses the clinical and cost barriers of dual-energy CT technology while maintaining diagnostic image quality across different contrast phases.
AINeutralarXiv – CS AI · May 276/10
🧠Researchers propose CAT (Cross-scale Aligned Transformer), a new GAN training method that addresses the cross-scale trajectory misalignment problem in multi-stage image generation. By adding consistency regularization between intermediate and final outputs, CAT achieves state-of-the-art results on ImageNet-256 with one-step inference, reaching FID-50K of 1.56 after just 60 training epochs.
AIBullisharXiv – CS AI · Mar 176/10
🧠Researchers introduce Contrastive Noise Optimization, a new method that improves diversity in text-to-image AI generation by optimizing initial noise patterns rather than intermediate outputs. The technique uses contrastive loss to maximize diversity while preserving image quality, achieving superior results across multiple text-to-image model architectures.
AIBullisharXiv – CS AI · Mar 55/10
🧠Researchers present Export3D, a new AI method for creating 3D-aware portrait animations from a single image with controllable facial expressions and camera angles. The technique uses a tri-plane generator and contrastive pre-training to avoid unwanted appearance changes when transferring expressions between different identities.