y0news
AnalyticsDigestsSourcesTopicsRSSAICrypto
🤖All92,136🧠AI22,940⛓️Crypto17,363💎DeFi1,799🤖AI × Crypto1,480📰General48,554
🧠

AI

22,940 AI articles curated from 50+ sources with AI-powered sentiment analysis, importance scoring, and key takeaways.

22940 articles
AIBullisharXiv – CS AI · Jun 236/10
🧠

Reference-Free Assessment of Physical Consistency in World Model-based Video Generation

Researchers introduced reference-free metrics for evaluating physical consistency in AI-generated videos, addressing a critical gap in world model evaluation. Using DROID-SLAM and SEA-RAFT technologies, the approach improved task success rates by over 8% and enables precise localization of physical artifacts, narrowing the simulation-to-reality gap for robotic applications.

AIBullisharXiv – CS AI · Jun 236/10
🧠

MetaPS: Adaptive Programmatic Strategy Selection for Market Agents

Researchers introduce MetaPS, a framework that enables AI agents to adaptively select from a library of pre-programmed trading strategies based on market conditions, rather than generating actions directly. The system uses market simulations to train models on when to deploy specific strategies, demonstrating consistent improvements across model sizes and outperforming fixed-strategy baselines and direct LLM decision-making approaches.

AINeutralarXiv – CS AI · Jun 236/10
🧠

PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems

Researchers introduced PlanBench-XL, a benchmark testing how LLM agents plan and execute tasks across 1,665 tools in realistic scenarios. The study reveals significant vulnerabilities in current AI systems, with performance dropping from 51.9% to 11.36% accuracy when tools fail or behave unexpectedly, exposing critical gaps in adaptive planning capabilities.

🧠 GPT-5
AINeutralarXiv – CS AI · Jun 236/10
🧠

Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent

Researchers evaluated whether structural codebase indexing improves coding agent performance by running controlled experiments with Claude Opus 4.7 across standardized benchmarks. Results show the index significantly improves code localization and task resolution rates without increasing costs, and outperforms simpler retrieval baselines, suggesting structural ranking becomes valuable for multi-file code changes.

🧠 Claude🧠 Opus
AINeutralarXiv – CS AI · Jun 236/10
🧠

SVGym (SciVerseGym): An Environment for Reinforcement Learning and Bayesian Optimization in Crystal Discovery

SVGym (SciVerseGym) is a new open-source framework that standardizes reinforcement learning workflows for automated crystal discovery by treating materials design as a Markov decision process. The environment decouples agent logic from materials infrastructure, enabling researchers to apply machine learning algorithms to accelerate the discovery of new materials with desired properties.

AIBullisharXiv – CS AI · Jun 236/10
🧠

Efficient Multimodal Clinical Question Answering for Pulmonary Embolism Risk Assessment

Researchers have developed a benchmark for evaluating efficient multimodal language models on pulmonary embolism diagnosis and risk assessment using a dataset of 23,248 CTPA studies. The study demonstrates that compact models like Gemma4 perform significantly better when combining imaging and electronic health record data, with diagnostic tasks outperforming prognostic predictions.

AINeutralarXiv – CS AI · Jun 236/10
🧠

Self-Evolving Cognitive Framework via Causal World Modeling for Embodied Scientific Intelligence

Researchers propose a self-evolving cognitive framework that moves embodied AI systems beyond predictive modeling toward causal reasoning and scientific intelligence. The approach integrates causal world modeling, intervention-driven reasoning, and continual refinement, enabling AI to learn through active experimentation rather than passive prediction.

AINeutralarXiv – CS AI · Jun 236/10
🧠

PRIME: Evaluating Prompt Resolution Under Incompatible Instructions in LLMs

Researchers introduce PRIME, a framework for evaluating how large language models handle conflicting instructions, revealing that conflict type significantly impacts model behavior regardless of scale. The study of five instruction-tuned LLMs exposes critical gaps in current benchmarking methods that assess instructions in isolation, demonstrating that real-world instruction-following capabilities cannot be accurately measured without testing competing directives.

AINeutralarXiv – CS AI · Jun 236/10
🧠

SCOPE: Evolving Symbolic World for Planning in Open-Ended Environments

Researchers introduce SCOPE, a self-adaptive framework that enhances Vision-Language Models' planning capabilities by refining symbolic representations of open-ended environments through iterative execution feedback. The system combines symbolic validation with adaptive memory mechanisms to improve long-horizon planning success rates and cross-task generalization in complex embodied AI scenarios.

AINeutralarXiv – CS AI · Jun 236/10
🧠

Deep Learning-Based Sign Language Recognition from Videos and Cross-Lingual Translation to Indian Vernaculars

Researchers have developed a deep learning pipeline that recognizes sign language gestures from videos and translates them into Indian languages using VideoMAE and Meta's NLLB-200 model. The system achieves 78% validation accuracy on a 13-class dataset and demonstrates practical accessibility applications, though it currently handles isolated words rather than continuous signing.

🏢 Meta
AINeutralarXiv – CS AI · Jun 236/10
🧠

Grounded Scaling: Why Agentic AI Needs Deterministic Environments

A new arXiv paper argues that agentic AI systems require deterministic environments to scale effectively, proposing that environment determinism is a critical binding constraint for AI progress alongside compute growth. The authors introduce a Supply Certainty Index and five-level Determinism Maturity Model to operationalize the framework for tasks with verifiable economic or physical outcomes.

AINeutralarXiv – CS AI · Jun 236/10
🧠

MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop

MacAgentBench introduces a comprehensive macOS agent benchmark with 676 tasks across 25 applications, enabling more rigorous evaluation of computer use agents (CUAs) like those deployed on Mac Mini. The study reveals that Claude Opus 4.6 on OpenClaw achieves 73.7% Pass@1, with skill libraries driving performance more than framework design, while fine-grained scoring exposes significant differences in sub-goal completion among models with similar overall scores.

🧠 Claude🧠 Opus
AINeutralarXiv – CS AI · Jun 236/10
🧠

Text2DSL: LLM-Based Code Generation for Domain-Specific Languages

Researchers introduce Text2DSL, a framework for automatically generating domain-specific language (DSL) code from natural language using large language models, validated on 4,204 Polkit security policy rules. The study demonstrates that providing structured context like BNF grammar and API specifications dramatically improves code generation accuracy to 98.6-99.4% syntactic validity across different model scales without requiring fine-tuning.

AINeutralarXiv – CS AI · Jun 236/10
🧠

SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment

SkillAudit introduces an automated framework for evaluating AI agent skills independently of fixed task benchmarks, addressing a critical gap in skill marketplaces. The research reveals that over 7% of real-world skill packages exhibit risky behavior, highlighting the need for systematic assessment tools as AI skill ecosystems expand.

AINeutralarXiv – CS AI · Jun 236/10
🧠

AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding Agent

Researchers introduce AgentLens, a white-box defense framework that detects and mitigates safety risks in multi-turn LLM coding agents by intervening in mechanistic subspaces. The framework achieves strong safety detection performance through step-level hidden representation analysis, addressing the limitations of external guardrails in capturing evolving execution risks.

AINeutralarXiv – CS AI · Jun 236/10
🧠

Safety-Aware Evaluation of LLM-Generated Driver Intervention Messages through Multi-Task Risk Fusion

Researchers propose the Driver Safety-Aware Intervention Score (DSAIS), a domain-specific metric for evaluating LLM-generated driver safety messages across five dimensions including risk-urgency alignment and cognitive load. The framework integrates multi-task recognition outputs through risk fusion and achieves strong inter-rater reliability (ICC 0.798-0.840), demonstrating that compact local LLMs outperform API-based models for in-vehicle deployment.

AIBullisharXiv – CS AI · Jun 236/10
🧠

Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation

Researchers introduce STREAM, a diffusion transformer model that generates danceable choreography from text and music by decoupling their conditioning pathways, preventing acoustic dominance from overwhelming semantic control. The team releases Motorica++, an enhanced dataset with semantic annotations, and proposes new evaluation metrics (Exchange Evaluation Protocol and Editable Dance Score) to measure zero-shot editability in generative motion synthesis.

AINeutralarXiv – CS AI · Jun 235/10
🧠

Learning Filters with Certainty

Researchers propose enhancing Counting Bloom Filters (CBFs) by leveraging certainty signals from hash collision information to improve machine learning model accuracy. This work demonstrates how traditional data structure design can be refined to provide probabilistic confidence metrics, enabling hybrid ML-filter architectures to make more informed decisions in applications like caching and anomaly detection.

AINeutralarXiv – CS AI · Jun 236/10
🧠

The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models

Researchers propose a comprehensive uncertainty quantification (UQ) framework for large language models, breaking down sources of error into input-level, parameter-level, token-level, and decoding-process components. Testing 21 UQ methods across Qwen3, Llama 3.2, and DeepSeek-V3 reveals that consensus-based approaches consistently outperform alternatives, while larger models exhibit lower uncertainty estimates according to an empirical scaling law.

🧠 Llama
AINeutralarXiv – CS AI · Jun 236/10
🧠

A Formula-Driven Survey and Research Agenda for On-Policy Distillation

This arXiv paper presents a comprehensive taxonomy and research framework for on-policy distillation (OPD), a technique for training large language models using feedback from current or recent student policies. The work moves beyond single loss functions to analyze OPD as a systematic feedback-to-update problem, introducing new methods like Counterfactual Routed OPD (CR-OPD) and identifying critical mechanisms affecting model stability and performance.

AINeutralarXiv – CS AI · Jun 235/10
🧠

AI-Assisted Help-Seeking Trajectories in Programming Education from an SRL-Informed Perspective

A study of 71 university students' interactions with generative AI in introductory Python programming reveals that most use AI reactively for troubleshooting rather than as a planned learning tool. While AI-assisted help-seeking patterns didn't significantly affect task scores, they substantially influenced the number of code submissions required, suggesting that how students engage with AI matters more than whether they use it.

AIBullisharXiv – CS AI · Jun 236/10
🧠

MINCE: Shrinking LLM Evaluation Datasets via Few-Model Monte Carlo Calibration

Researchers introduce MINCE, a novel method that significantly reduces the computational cost of evaluating large language models by intelligently shrinking benchmark datasets. Using Monte Carlo simulation with minimal calibration models, MINCE achieves 54-89% dataset size reductions while maintaining accuracy within acceptable drift thresholds, enabling 2.7-8.1x faster GPU evaluations.

AIBullisharXiv – CS AI · Jun 236/10
🧠

RaMem: Contextual Reinstatement for Long-term Agentic Memory

Researchers introduce RaMem, a framework that solves the 'context collapse' problem in long-term LLM agent memory systems by recontextualizing retrieved memory fragments with their original episodic conditions. The approach uses evidence anchoring, condition induction, validity-aware retrieval, and context-preserved synthesis to improve memory relevance verification, achieving over 10% F1 improvement across benchmarks.

AINeutralarXiv – CS AI · Jun 236/10
🧠

AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions

Researchers propose that agentic AI systems are transitioning from computational tools into autonomous "AI scientists" capable of accelerating scientific discovery across literature synthesis, hypothesis generation, and model verification. The paper argues this requires fundamental institutional reforms around verification, accountability, and safety, and introduces Denario as a prototype multi-agent framework that can explore hypothesis spaces beyond human capability.

AIBullisharXiv – CS AI · Jun 236/10
🧠

Agent-as-a-Router: Agentic Model Routing for Coding Tasks

Researchers propose Agent-as-a-Router, a framework that dynamically routes coding tasks to the most suitable LLM among multiple providers by accumulating execution-grounded experience during deployment. The approach, instantiated as ACRouter, demonstrates 15.3% performance gains over static routers and introduces CodeRouterBench, a benchmark with ~10K tasks from 8 frontier LLMs, addressing the critical need for intelligent model selection in multi-provider environments.

← PrevPage 272 of 918Next →
Filters
Sentiment
Importance
Sort
Stay Updated
Everything combined