y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto
🤖All86,127🧠AI22,940⛓️Crypto17,361💎DeFi1,798🤖AI × Crypto1,480📰General42,548
🧠

AI

22,940 AI articles curated from 50+ sources with AI-powered sentiment analysis, importance scoring, and key takeaways.

22940 articles
AIBullisharXiv – CS AI · Jun 97/10
🧠

SLMJury: Can Small Language Models Judge as Well as Large Ones?

Researchers introduce SLMJury, a framework demonstrating that small language models (0.6B-14B parameters) can match or exceed large language models as judges for evaluating AI outputs. The study reveals that model size alone doesn't determine judging capability, with performance varying significantly by task domain and judgment type, challenging assumptions about requiring expensive proprietary LLMs for automated evaluation.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Beyond Accuracy: Interpreting Topic Representation in Suicide Ideation Detection Models

Researchers demonstrate that suicide ideation detection models trained with topic-augmented datasets develop more interpretable internal representations of psychological risk factors. The study moves beyond standard accuracy metrics to examine how AI systems encode mental health concepts, revealing that augmentation clarifies underrepresented factors like immigration stress, family issues, and financial crisis.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Complement or substitute? How AI increases the demand for human skills

A comprehensive empirical study analyzing 30 million US, UK, and Australian job postings finds that AI adoption increases demand for complementary human skills like analytical thinking and resilience rather than simply replacing workers. The research reveals significant wage premiums for these soft skills in AI-adjacent roles and spillover effects where AI diffusion reduces demand for substitutable tasks across entire industries and regions.

AINeutralarXiv – CS AI · Jun 97/10
🧠

SENTRY: Statistical Reliability Analysis of Vision Transformers Under Soft Errors

Researchers present SENTRY, a statistical fault injection framework that efficiently evaluates Vision Transformers' reliability against soft errors in safety-critical applications. The method achieves formal reliability guarantees using finite-population sampling theory, reducing experimental costs by up to 10,700x while identifying critical vulnerabilities in normalization layers and IEEE-754 exponent bits.

AIBullisharXiv – CS AI · Jun 97/10
🧠

ScaleSweep: Accurate NVFP4 Post-Training Quantization of LLMs via Block Scale Initialization

ScaleSweep introduces an optimized block scale initialization method for NVFP4 quantization of large language models, improving upon traditional AbsMax approaches. The technique theoretically bounds the search space and empirically achieves 93% performance retention under aggressive 4-bit quantization, advancing hardware-efficient AI inference.

🧠 Llama
AINeutralarXiv – CS AI · Jun 97/10
🧠

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence

A position paper argues that Anthropomorphic Misalignment Research (AMR) studies often lack sufficient empirical rigor to support critical AI safety decisions. The authors propose an evidence framework and diagnostic checklist to strengthen methodological standards and ensure AI risk claims rest on solid foundations.

AIBearisharXiv – CS AI · Jun 97/10
🧠

Can You Trust What You See? Human and AI Detection of Synthetic Legal Evidence

Researchers evaluated humans and advanced AI models on detecting synthetic legal evidence, finding both groups unreliable authenticators. Human accuracy dropped to near-chance levels (48-51%) against leading image generators, while AI models achieved perfect specificity but missed most synthetic outputs, suggesting visual evidence requires multi-layered verification in legal proceedings.

🧠 GPT-5🧠 Gemini
AIBearisharXiv – CS AI · Jun 97/10
🧠

VisualLeakBench: Reproducible Action-Boundary Propagation Failures in Vision-Language Agents

Researchers introduce VisualLeakBench, a 500-image benchmark that reveals critical security vulnerabilities in vision-language agents, where sensitive information visible in screenshots and documents is propagated into tool arguments. Testing four production VLM systems shows baseline failure rates of 78.8% for personally identifiable information and 85.5% for unsafe text, with defensive prompts reducing PII propagation but leaving unsafe-text leakage at 52.6%.

AINeutralarXiv – CS AI · Jun 97/10
🧠

Performative Learning Theory

Researchers present a theoretical framework analyzing how predictive models that influence real-world outcomes affect generalization and learning capacity. The study reveals a fundamental trade-off: models that significantly impact data generate less reliable insights about future populations, with implications for algorithmic systems in employment, finance, and other consequential domains.

AIBullisharXiv – CS AI · Jun 97/10
🧠

FineGen: A VLM-based Multi-Agent Framework for Fine-Grained Image-Text Dataset Construction

FineGen is a VLM-based multi-agent framework that automatically constructs vision-language datasets by generating hard negative samples through a Generation-Verification-Correction pipeline. The resulting FineGen-100K dataset contains 147,000+ attribute-specific hard negatives and demonstrates a 14.4% accuracy improvement on fine-grained object detection benchmarks, addressing a critical gap in existing datasets.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation

Researchers introduce Item Response Scaling Laws (IRSL), a framework that dramatically reduces computational costs for estimating language model performance by decomposing the problem into model ability and question difficulty components. The approach achieves 99.9% reduction in required evaluation samples while maintaining or exceeding accuracy of traditional scaling law methods.

AINeutralarXiv – CS AI · Jun 97/10
🧠

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

Researchers introduced ResearchClawBench, a comprehensive benchmark with 40 tasks across 10 scientific domains designed to evaluate AI agents' ability to conduct autonomous scientific research. Current leading systems like Claude Code and Claude-Opus-4 score only 20-21.5 points, revealing significant gaps in experimental design, evidence synthesis, and scientific reasoning capabilities.

🧠 Claude
AIBullisharXiv – CS AI · Jun 97/10
🧠

Liberating LLM Capabilities in Full-Duplex Speech Models

Researchers introduce Listen-Write-Speak (LWS), a new paradigm for speech-based large language models that enables simultaneous text output alongside spoken responses. The approach leverages a single autoregressive LLM with a Token Schema to unlock text-native capabilities like code generation and structured analysis in real-time conversational AI without architectural modifications.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling

Researchers present a production-deployed recommendation system that scales short-form video suggestions to billion-user scale by replacing traditional Video IDs with semantic-native representations and introducing a compression transformer to reduce computational complexity. The framework achieves order-of-magnitude improvements in memory efficiency and enables longer user behavior sequences, delivering measurable gains in user engagement and content consumption metrics.

AIBullisharXiv – CS AI · Jun 97/10
🧠

AgentCompile: An LLM-Guided Compiler for Direct CUDA Inference

AgentCompile is an LLM-guided CUDA inference compiler that uses large language models to optimize transformer model execution on GPUs. The system achieves 4-5.66x speedup over PyTorch across popular models like Qwen and Llama through intelligent specialization decisions and empirical validation.

🧠 Llama
AIBullisharXiv – CS AI · Jun 97/10
🧠

SurfDesign: Effective Protein Design on Molecular Surfaces

Researchers introduce SurfDesign, a novel protein design framework that conditions on molecular surface geometry rather than just backbone structure, integrating surface-based equivariant message passing with pretrained protein language models. The method significantly outperforms existing approaches on de novo binder and enzyme design benchmarks, demonstrating that manifold-aware surface representations provide a more effective foundation for functional protein design.

AIBullisharXiv – CS AI · Jun 97/10
🧠

DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home

Researchers introduce DIYHealth Suite, a comprehensive framework including a 900K-sample multimodal dataset, adaptive foundation model, and benchmark for home-based health management powered by generative AI. The framework addresses critical gaps in making healthcare accessible outside clinical settings through standardized tools for diverse home care scenarios.

AIBearisharXiv – CS AI · Jun 97/10
🧠

Beware of GeeksBearing Gifts: Building True EU Frontier AI Sovereignty

EU researchers propose a comprehensive framework for achieving frontier AI sovereignty across five pillars and five technology stack layers, addressing Europe's structural dependence on US and Chinese AI models. The analysis reveals fragmentation in existing EU policy and demonstrates how the proposed sovereignty-centered approach could guide strategic interventions across 92 Commission initiatives.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Dynamic Distributed Constraint Optimization and Metareasoning for Continual, Large-Scale Satellite Operations

Researchers have developed a novel framework for autonomously scheduling observations across large satellite constellations using distributed constraint optimization. The work introduces the dynamic multi-satellite constellation observation scheduling problem (DCOSP) and the D-NSS algorithm, which enables satellites to coordinate efficiently with minimal communication overhead—a critical advancement for NASA's FAME mission demonstrating distributed multi-agent AI in space.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Enabling KV Caching of Shared Prefix for Diffusion Language Models

Researchers introduce bicache, a novel KV caching technique that enables efficient serving of diffusion language models (DLMs) with shared prefixes. Unlike traditional LLMs, DLMs use bidirectional attention, which invalidates conventional caching methods and causes accuracy collapse. Bicache dynamically identifies safe layer depths for prefix reuse, achieving 36-98% throughput improvements.

AIBullisharXiv – CS AI · Jun 97/10
🧠

BRAIN: Bayesian Reasoning via Active Inference for Agentic and Embodied Intelligence in Mobile Networks

Researchers propose BRAIN, a Bayesian reasoning AI agent for 6G mobile networks that uses active inference to improve decision-making transparency and adaptability. Unlike conventional deep reinforcement learning approaches, BRAIN demonstrates 28.3% better robustness to traffic shifts without retraining and provides human-interpretable explanations of its network resource allocation decisions.

AIBullisharXiv – CS AI · Jun 97/10
🧠

PRISM: Recovering Instruction Sets from Language Model Activations

Researchers introduce PRISM, a new AI system that decodes hidden states from language models to reveal the complete set of active instructions guiding their behavior. This advancement addresses a critical security gap in monitoring deployed LLM agents by detecting unintended objectives, prompt injections, and hidden constraints that models may follow without explicit output indication.

AINeutralarXiv – CS AI · Jun 97/10
🧠

AI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset Valuation

Researchers demonstrate that AI agents' performance in drug-asset valuation is fundamentally limited by access to proprietary data rather than reasoning quality alone. A three-arm experiment shows that adding reasoning scaffolds and structured tools improves calibration but cannot overcome gaps in underlying evidence, with proprietary datasets enabling 96% recovery of expert valuations versus 38% for public-data-only systems.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Next-Token Prediction Learns Generalisable Representations of Sleep Physiology

Researchers introduce Hypnos, a multi-modal foundation model trained on next-token prediction that learns generalizable representations of sleep physiology from over 20,000 polysomnography recordings across eight sensing modalities. The model achieves performance parity with supervised baselines on sleep stage classification while using 100× less labeled data and demonstrates cross-domain generalization by outperforming specialized models on daytime cardiac tasks.

AIBullisharXiv – CS AI · Jun 97/10
🧠

Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text

Researchers propose optical reasoning, a novel approach that uses images as the primary medium for AI reasoning tasks rather than text. The method demonstrates 28.57% token reduction on language tasks and 16% on multimodal tasks while matching or exceeding traditional text-based reasoning performance across mathematical, scientific, and multimodal benchmarks.

← PrevPage 51 of 918Next →
Filters
Sentiment
Importance
Sort
Stay Updated
Everything combined