y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto
🤖All95,013🧠AI22,940⛓️Crypto17,363💎DeFi1,799🤖AI × Crypto1,480📰General51,431

AI × Crypto News Feed

Real-time AI-curated news from 95,027+ articles across 50+ sources. Sentiment analysis, importance scoring, and key takeaways — updated every 15 minutes.

95027 articles
AIBullisharXiv – CS AI · Jun 97/10
🧠

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Researchers introduce TAME, a trust-aware memory evolution framework that addresses the vulnerability of AI agents to safety misalignment during test-time learning. The system uses paired Executor and Evaluator components to selectively reinforce and reuse agent memories, demonstrating 14.6 percentage point accuracy improvements on mathematical benchmarks while maintaining trustworthiness.

🧠 GPT-5
AIBullisharXiv – CS AI · Jun 97/10
🧠

ConflictRAG: Detecting and Resolving Knowledge Conflicts in Retrieval Augmented Generation

ConflictRAG introduces a novel framework for detecting and resolving contradictory information in Retrieval-Augmented Generation systems, achieving 88.7% conflict-detection accuracy while reducing API costs by 62%. The system combines cost-efficient embedding-based detection with selective LLM refinement and demonstrates 5.3-6.1% improvements in answer correctness across multiple benchmarks.

AIBullisharXiv – CS AI · Jun 97/10
🧠

MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting

Researchers propose MMR-GRPO, a training optimization technique that accelerates Group Relative Policy Optimization (GRPO) for mathematical reasoning models by reweighting rewards based on completion diversity. The method achieves comparable performance while reducing training time by 70.2% and training steps by 47.9%, demonstrating consistent improvements across multiple model sizes and benchmarks.

AIBullisharXiv – CS AI · Jun 97/10
🧠

MedVision: Benchmarking Quantitative Medical Image Analysis

Researchers introduce MedVision, a large-scale benchmark dataset with 30.8 million image-annotation pairs designed to evaluate and improve vision-language models (VLMs) on quantitative medical image analysis tasks. The work demonstrates that current VLMs perform poorly on clinical quantitative reasoning—such as tumor measurement and joint angle assessment—but can be significantly improved through supervised and reinforcement fine-tuning.

AIBullisharXiv – CS AI · Jun 97/10
🧠

A large-scale nanocrystal database with aligned synthesis and properties enabling generative inverse design

Researchers have created a large-scale database of 160,000 aligned nanocrystal synthesis-property entries using AI, enabling generative inverse design for materials discovery. The system successfully predicts viable synthesis routes for both established and novel nanocrystals, including counter-intuitive formulations validated experimentally, demonstrating AI's potential to accelerate materials science beyond traditional trial-and-error methods.

AIBearisharXiv – CS AI · Jun 97/10
🧠

Multimodal Generative Engine Optimization: Rank Manipulation for Vision-Language Model Rankers

Researchers demonstrate a critical vulnerability in Vision-Language Models (VLMs) used for ranking and recommendation systems through Multimodal Generative Engine Optimization (MGEO), showing that adversaries can manipulate ranking decisions by combining imperceptible image perturbations with crafted text. This attack exploits the deep cross-modal knowledge coupling within VLMs, revealing fundamental weaknesses in how these models ground and apply multimodal information.

AIBullisharXiv – CS AI · Jun 97/10
🧠

BCG-FM: A Foundation Model for Ambient Cardiac Health Sensing

Researchers introduce BCG-FM, a foundation model trained on 2.75 million hours of ballistocardiography data from nearly 146,000 individuals, enabling non-invasive cardiac health monitoring through piezoelectric bed sensors. The model achieves state-of-the-art biological age estimation and demonstrates clinical relevance across multiple health conditions without requiring deliberate user action.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT

Researchers introduce OptiKIT, an open-source distributed framework that automates LLM optimization for enterprise deployments, delivering over 2x GPU throughput improvements while eliminating the need for specialized optimization expertise. The system democratizes model compression and tuning through dynamic resource allocation and intelligent pipeline orchestration, addressing a critical bottleneck in scaling AI initiatives within compute-constrained environments.

AIBullisharXiv – CS AI · Jun 97/10
🧠

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation

CURE is a curriculum learning framework that improves medical vision-language models' ability to generate accurate radiology reports with better visual grounding. The method achieves significant gains in grounding accuracy (+0.35 IoU), report quality (+0.192 CXRFEScore), and hallucination reduction (18.6%) without requiring additional training data.

🏢 Hugging Face
AINeutralarXiv – CS AI · Jun 97/10
🧠

Performative Learning Theory

Researchers present a theoretical framework analyzing how predictive models that influence real-world outcomes affect generalization and learning capacity. The study reveals a fundamental trade-off: models that significantly impact data generate less reliable insights about future populations, with implications for algorithmic systems in employment, finance, and other consequential domains.

AIBullisharXiv – CS AI · Jun 97/10
🧠

More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD)

Researchers propose Reset-and-Discard (ReD), a novel querying method that improves large language model inference efficiency by optimizing the coverage@cost metric—the number of unique questions answered within a fixed budget. The technique reduces computational attempts, tokens, and financial costs needed to achieve desired performance levels across coding, math, and reasoning tasks.

AIBullisharXiv – CS AI · Jun 97/10
🧠

MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering

Researchers introduce MEnvAgent, a framework for automatically constructing executable software environments across multiple programming languages, addressing a critical bottleneck in LLM agent training. The system generates verifiable datasets and reduces computational costs by 43%, enabling the creation of MEnvData-SWE, the largest open-source polyglot dataset of Docker environments for software engineering tasks.

AIBearisharXiv – CS AI · Jun 97/10
🧠

Pretrained, Frozen, Still Leaking: Auditing Cross-Encoder Attribute Transfer in EEG Foundation Models

Researchers demonstrate that popular EEG foundation models (BIOT, LaBraM, EEGPT) leak sensitive neurological attributes despite appearing secure under individual audits. A cross-encoder transfer attack shows that attribute decoders trained on one frozen model successfully transfer to others, indicating shared vulnerabilities that standard defenses like differential privacy fail to adequately address.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation

Researchers introduce Item Response Scaling Laws (IRSL), a framework that dramatically reduces computational costs for estimating language model performance by decomposing the problem into model ability and question difficulty components. The approach achieves 99.9% reduction in required evaluation samples while maintaining or exceeding accuracy of traditional scaling law methods.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Single-Cell Cross-Modal Transfer by Adversarial Fine-Tuning of Foundation Models

Researchers propose a foundation model approach using adversarial fine-tuning to translate between unpaired spatial transcriptomics and single-cell RNA sequencing data. The method addresses the scarcity of paired datasets by leveraging the abundance of individual modalities, outperforming existing multi-omics translation approaches.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Beyond Accuracy: Interpreting Topic Representation in Suicide Ideation Detection Models

Researchers demonstrate that suicide ideation detection models trained with topic-augmented datasets develop more interpretable internal representations of psychological risk factors. The study moves beyond standard accuracy metrics to examine how AI systems encode mental health concepts, revealing that augmentation clarifies underrepresented factors like immigration stress, family issues, and financial crisis.

AINeutralarXiv – CS AI · Jun 97/10
🧠

SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

Researchers introduce SWE-Marathon, a benchmark testing AI agents on 20 ultra-long-horizon software engineering tasks requiring millions of tokens and hours of sustained work. Current frontier coding agents solve fewer than 30% of tasks, revealing critical gaps in planning, self-verification, and memory management that limit real-world deployment.

AINeutralarXiv – CS AI · Jun 97/10
🧠

Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units

Researchers introduce Mechanistic Data Attribution (MDA), a framework using Influence Functions to trace interpretable units in large language models back to specific training samples. Through experiments on Pythia models, they demonstrate that targeted removal or augmentation of high-influence training samples causally affects the emergence of interpretable circuits, while providing direct evidence linking induction heads to in-context learning capabilities.

AIBearisharXiv – CS AI · Jun 97/10
🧠

MLingualFC: Evaluating Jailbreak Vulnerabilities in Multilingual Vision-Language Models

Researchers introduced MLingualFC, a benchmark revealing significant safety vulnerabilities in multilingual Vision-Language Models through flowchart-based jailbreak attacks across five languages. The study demonstrates that current VLM safety mechanisms fail to generalize across linguistic and visual modalities, with Latin script languages showing substantially higher attack success rates than non-Latin scripts like Punjabi.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

Researchers present RTPurbo, a method that transforms standard full-attention language models into efficient sparse models within just hundreds of training steps. By leveraging the observation that LLMs are intrinsically sparse, the approach achieves up to 9.36× speedup during prefill and 2.01× during decode at 1M context length while maintaining near-lossless accuracy.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design

Meta researchers have developed Kunlun, a scalable architecture for recommendation systems that establishes predictable scaling laws by improving model efficiency from 17% to 37% on GPU utilization. The system combines low-level optimizations like Generalized Dot-Product Attention with high-level innovations to double scaling efficiency, now deployed across Meta's advertising infrastructure.

🏢 Nvidia
AIBearisharXiv – CS AI · Jun 97/10
🧠

Can You Trust What You See? Human and AI Detection of Synthetic Legal Evidence

Researchers evaluated humans and advanced AI models on detecting synthetic legal evidence, finding both groups unreliable authenticators. Human accuracy dropped to near-chance levels (48-51%) against leading image generators, while AI models achieved perfect specificity but missed most synthetic outputs, suggesting visual evidence requires multi-layered verification in legal proceedings.

🧠 GPT-5🧠 Gemini
AIBullisharXiv – CS AI · Jun 97/10
🧠

AgentCompile: An LLM-Guided Compiler for Direct CUDA Inference

AgentCompile is an LLM-guided CUDA inference compiler that uses large language models to optimize transformer model execution on GPUs. The system achieves 4-5.66x speedup over PyTorch across popular models like Qwen and Llama through intelligent specialization decisions and empirical validation.

🧠 Llama
AIBullisharXiv – CS AI · Jun 97/10
🧠

Post-Trained MoE Can Skip Half Experts via Self-Distillation

Researchers introduced ZEDA, a framework that converts fully-trained Mixture-of-Experts language models into dynamic variants capable of skipping unnecessary experts, reducing computational requirements by over 50% with minimal accuracy loss. The method uses self-distillation to adapt post-trained models without retraining from scratch, achieving ~1.20x end-to-end inference speedup on major language models.

AIBearisharXiv – CS AI · Jun 97/10
🧠

Adversarial Robustness of Activation Steering in Large Language Models

Researchers demonstrate that activation steering, a popular training-free method for controlling large language model behavior, is highly vulnerable to adversarial text perturbations. The study reveals that attacks can degrade steering effectiveness by up to 64% and cause optimal layer selections to shift by 17 positions, exposing structural brittleness that poses risks for real-world deployment.

🏢 Anthropic
← PrevPage 223 of 3802Next →
Filters
Sentiment
Importance
Sort
Stay Updated
Everything combined